LLMs are everywhere. We use them to generate code, summarize documents and answer questions throughout the day. For many people, this is what using AI now looks like.
That success can make it tempting to approach every AI problem through language. In applications such as robotics, drone navigation, and automotive driver assistance, systems need to predict how their surroundings will change when they act. They also operate with limited power, memory, and time to make decisions.
Could we predict the consequences of actions directly on the device, without relying on a datacenter for each decision? Compact world models make that an interesting research direction.
Where LLMs reach their limits
Current frontier models such as GPT-6 Astra, Claude Fable 5.1, and Kimi K3 work across text, images, code, and software tools. These capabilities help interpret goals and organize tasks. Physical control adds another requirement: predicting how a particular machine and its surroundings will respond to an action. Success in coding or visual reasoning does not by itself establish a reliable model of those dynamics.
Hardware adds another constraint. Kimi K3 has 2.8 trillion total parameters, with only a subset active for each token. Its developers recommend deployment on systems with 64 or more accelerators. That recommendation illustrates the scale of frontier inference: datacenter infrastructure far beyond a typical embedded device's power and memory budget.
One option is to run inference in a remote datacenter and exchange observations and predictions over the network. That reduces onboard computation, but adds latency and availability risks. For physical control with tight deadlines, unpredictable delays, disconnections, or service outages can be unacceptable: the device must still respond when the remote system is unavailable.
For physical autonomy, both the model's objective and its hardware footprint matter. We need to study models that learn from observations and interaction, predict the consequences of actions, and fit the available computing budget.
World models
A world model learns a representation of an environment and how it changes. For control, it predicts future states from observations and proposed actions. Repeating that prediction lets a system compare sequences of actions before executing them.
These predictions can support planning during operation or provide imagined experience for training a policy, the rule that selects actions. Both approaches can reduce the amount of trial and error needed in the physical environment. The same learned dynamics can support different goals within that environment.
This is why world models matter for physical AI: machines that perceive, decide, and act in the world. Useful applications include navigation, manipulation, collision anticipation, and adapting behavior when conditions change. The model needs enough understanding of the relevant environment to guide those decisions.
For example, World Labs' Atlas reconstructs and simulates environments from images and video. World Labs shows it producing 3D scenes and generating the images and depth observations that simulated robots would see. These capabilities can help build varied training and evaluation environments for navigation and manipulation.
Empiric Earth applies world models to road safety. Its BADAS 2.0 model processes vehicle camera footage with a V-JEPA2-based architecture to predict how a scene may evolve and estimate collision risk. Those predictions can support driver assistance and fleet safety systems.
For autonomous flight, SkyDreamer learns drone control with a world model. Its encoder, sequence model, and action policy run onboard an NVIDIA Jetson Orin NX.
LeWorldModel
LeWorldModel, or LeWM, is a compact example of this research. Its March 2026 paper is by Lucas Maes, Quentin Le Lidec, Damien Scieur, Yann LeCun, and Randall Balestriero.
LeCun's involvement gives the work considerable scientific weight. He shared the 2018 ACM A.M. Turing Award with Yoshua Bengio and Geoffrey Hinton for their contributions to deep learning. He is a coauthor actively developing this direction with the research team.
LeWM follows the Joint Embedding Predictive Architecture (JEPA) approach. It encodes images into compact representations, then predicts future representations conditioned on actions. Learning useful features lets the model avoid reconstructing every pixel of every possible future.
Its contribution is a simpler way to train the encoder and predictor together from recorded images and actions. Two training objectives combine accurate prediction with protection against representation collapse, where different observations become indistinguishable inside the model.
The authors describe roughly 15 million parameters, trainable on a single GPU, and evaluate control in simulated environments. That scale makes the released implementation accessible for research into efficient physical prediction.
Why the edge matters
Compact world models open a path to running prediction and planning at the edge, on the device collecting observations and taking action. Local execution avoids a network round trip for each decision and allows operation without a continuous cloud connection. Training can still happen on more powerful hardware.
For us, the research opportunity also extends to hardware. A custom FPGA implementation could tailor arithmetic and dataflow to the model and application, potentially delivering lower latency and power consumption than a platform such as NVIDIA Jetson. LeWorldModel gives us a concrete starting point to test that possibility against an optimized Jetson implementation, measuring the complete control loop at comparable prediction quality.
If you're working on world models or exploring how to implement them in hardware, reach out to us. We'd love to hear about your project and explore how we could work together to help you reach your goals.