Capek 0.5: Execution-Centric Vision-Language Model Organizes Embodied AI Capabilities by Function

New embodied AI model organizes vision-language capabilities around execution roles rather than isolated tasks, using specialist training and merging.

Capek 0.5: Execution-Centric Vision-Language Model Organizes Embodied AI Capabilities by Function

According to arxiv.org, researchers have introduced Capek 0.5, an embodied vision-language model that organizes AI capabilities around their functional roles during robot execution rather than by datasets or tasks. The model addresses a key challenge: robot execution is “inherently iterative,” with each action reshaping the scene and requiring continuous perception, reasoning, and verification.

The system structures embodied capabilities into four families based on their execution roles: Spatial Reasoning, Temporal Understanding, Action Guidance, and State Verification, according to the paper. Each capability is first developed by a dedicated specialist through reinforcement learning with verifiable rewards from a shared backbone, then consolidated into a single model through weight-space merging followed by routed policy-space distillation.

Capek 0.5 is available at 2B and 35B-A3B scales, the paper states. The researchers evaluated it using comprehensive benchmark suites including Capek-StateBench (a new state verification benchmark), a controlled study of capability retention, and closed-loop evaluation in simulated embodied environments.

According to arxiv.org, results show that Capek 0.5 “improves the large majority of matched benchmark rows over its initialization, retains all four specialized capabilities in one checkpoint with quantified losses, and transfers to closed-loop embodied task execution.”

The research addresses how existing approaches “typically develop these capabilities against isolated, task-specific objectives, leaving open how they should be organized and integrated around execution as a whole,” the paper notes.