共有:
Strategic Pivot

Multimodal pioneer Reka
steers toward robot minds

Reka's focus had been text, image, and video multimodal models. Now it is shifting its center of gravity beyond the screen — toward the physical world where robots move.

AI Navigate Editorial2026.07.257 min read

MULTIMODAL TEXT・IMAGE・VIDEO PIVOT WORLD LANGUAGE ACTION MODEL ROBOTICS・EMBODIED AI
01
Why Now

From chat to robots,
a quiet turn

While most AI labs have fought their battles in chatbots, Reka built its name on multimodal models spanning text, image, and video. Now the company says it is pursuing a "World Language Action Model," a full pivot toward physical-world foundation models aimed at robotics and embodied AI. It's a break from the screen-bound multimodal work that defined Reka until now. The company lays out its positioning on its official site.

This pivot doesn't look like an isolated move. Over the past few years, frontier labs have gradually shifted their center of gravity from AI that generates words to AI that can act in the physical world, and the best-funded labs have been widening their bets on robotics foundation models. Moving from multimodal models closed within a digital space of text, image, and video toward foundation models that let embodied robots perceive and act in the real world — Reka's strategic shift looks like one instance of that broader industry tremor.

Reka beforeReka now
Multimodal models across text, image, and videoFoundation-model research under a "World Language Action Model" banner
Focused on chat and content generationAimed at robotics and embodied AI
Everything resolved on a screenThe target is the physical world itself

From AI that only produces words,
to AI that acts in the world.


02
World Language Action Model

A three-layer bet,
spelled out in the name

The sequence "World," "Language," "Action" hints at a single model trained to unify perception, language, and action.

PERCEIVE GROUND IN LANGUAGE ACT Perceive → ground → act, looped to adapt to the environment (inferred)
FIG. One reading of the three-layer structure implied by the name "World Language Action Model" (an inferred diagram, not an official Reka illustration)

The name "World Language Action Model" itself signals that Reka is moving a step beyond its earlier "see, read, describe" posture. As far as the name reveals, the framing the company has adopted appears to aim at training a single model that captures the state of the physical world (World), represents it in language (Language), and outputs it as actual action (Action) — all in one unified system. That goes further than conventional multimodal models, which mainly learn correspondences between text and images, reaching instead into continuous action outputs like a robot's joint angles or grasping motions. That is a marked departure from Reka's earlier research. Still, no concrete training methodology or architectural detail has been disclosed yet, so this reading is necessarily inferred from the name and the stated direction rather than confirmed specifics.

03
Who It Matters To

Who this affects,
and how

The impact splits sharply by role.

Developers & Engineers

There is no indication yet of any general-availability robotics API or SDK — this remains a stated research direction. Anyone using Reka's existing chat or multimodal API is unlikely to see changes for now, but engineers who follow robotics are worth keeping an eye on Reka's future technical posts and papers.

PMs & Business Development

This is a space worth watching for future robotics partnerships or hardware tie-ins. Worth watching if you follow robotics — this pivot is exactly the kind of shift worth tracking against competing labs' moves, so you can gauge any ripple effects on your own product early.

Chatbot Users

No effect if you use Reka as a chat model. There's no indication that existing product availability or usability changes because of this shift, so this category can safely stay on the sidelines.

04
What's Next & Risks

What happens next,
and the hurdles ahead

The near-term focus is what technical detail or benchmarks Reka publishes under this new direction. For now this remains a stated direction, without a disclosed robot platform, partner, or performance figures. Practical next steps for anyone tracking this: watch Reka's official announcements and blog for model or evaluation-metric disclosures; if you rely on Reka's existing chat or multimodal API, confirm there's no near-term roadmap change; and weigh Reka's position against the broader robotics-foundation-model field — both major labs and dedicated startups already active there.

Optimism should be tempered, though. Building foundation models that act in the physical world is a notoriously long, capital-intensive process spanning data collection, simulation, and real-hardware validation. Well-funded major labs and dedicated robotics startups are already established in this space, and it's an open question how far a comparatively smaller lab like Reka can compete on the same field. Nothing in this announcement names a concrete product, benchmark, or timeline — it remains, for now, a statement of research direction.

Source: Reka · AI Navigate — Daily Update · 2026.07.25