SEPTEMBER 2026 · RESEARCH CONCEPT · v4
Physical Firmware
A structured world model for reusable physical intelligence.
A robot should not need to relearn the reusable structure of physical interaction whenever its task, objects, or body change. Physical Firmware asks whether a learned model of action consequences transfers better when selected geometry, mechanics, contact constraints, and uncertainty are explicit parts of its architecture.
Not an experimentally validated system. No Physical Firmware performance results are claimed here.
01 / A physical prediction layer
↳ World observations return to the belief.
↳ Planner proposals return to the predictor.
01 / THE PROBLEMKnowing what to do is not the same as predicting what will happen.
A robot pushes a part toward a fixture. Change the surface finish, the hidden weight, or the controller delay. The instruction stays the same. The consequences do not.
A policy maps observations and goals to actions. A world model predicts consequences, usually conditioned on candidate actions. A vision-language-action model, or VLA, is a policy that connects visual observations and language to robot actions. Modern VLAs already contain substantial implicit knowledge of physical interaction; generalist behavior and cross-robot transfer are demonstrated research directions, not capabilities this proposal invents.[1][7]
What remains difficult is learning reliably from interventions in an unforgiving world. Physical trials consume robot time, resets, operator attention, and sometimes hardware. Rare contact failures matter even when average performance is good. A camera may not reveal mass, center of mass, friction, or compliance. Small geometry errors can change a collision into a near miss; small trajectory errors can compound until a plan enters an unfamiliar state.
Contact also changes the kind of motion that is possible. Before touching a surface, a gripper moves freely. After touching, the same command may produce force, deformation, sticking, or slip. A policy changes the distribution of states it encounters as it improves—or as it makes a mistake. Prediction evaluated only on recorded demonstrations can therefore misrepresent prediction under the policy's own actions.
Learning behavior means learning useful choices. Learning dynamics means learning how interventions change the world. The two can share representations and training objectives. The unresolved question is whether selected explicit physical structure is a better inductive bias—a restriction or preference built into learning—for transfer, prediction, and control.
Some knowledge of physical interaction may survive changes of task, object, and embodiment more reliably when its predictive model is structured around entities, geometry, symmetries, constraints, hybrid events, and uncertainty.
02 / THE HYPOTHESISTrain broadly. Adapt lightly. Learn specifically.
Physical Firmware is a reusable, action-conditioned, probabilistic world model for physical interaction, with architectural structure derived from mechanics and learned residual dynamics derived from data. “Action-conditioned” means predictions depend on what the robot might do. “Probabilistic” means the output represents multiple plausible outcomes, not just one guessed trajectory. A residual is a learned correction for effects the chosen mechanics does not explain.
The computer analogy—hardware → firmware → operating system → application—is about reuse, not literal software placement. A physical prediction layer would be queried by task intelligence and grounded by robot-specific interfaces. It need not run on a microcontroller, live below the operating system, or own the control loop.
02 / What might transfer?
Shared physical core
Interaction operators; geometric structure; appropriate conservation and dissipation; contact mechanisms; predictive uncertainty.
Embodiment grounding
Kinematics, morphology, sensors, actuator dynamics, compliance, gripper geometry, latency, and robot-specific residuals.
Task intelligence
Instruction meaning, goals, costs, sequencing, behavior selection, VLA reasoning, and task policies.
Shared core + grounded embodiment + task objective → a usable robot system.
The intended benefit is to learn “how to achieve this outcome” while reusing knowledge of consequences. But “train once” would be misleading. New material classes, sensor configurations, or mechanisms may require core updates. Even within rigid manipulation, an adapter might become so large that the factorization stops being useful.
What is architectural, and what is learned?
| Candidate structure | Learned or inferred quantities | What it does not guarantee |
|---|---|---|
| Object graph and shared interactions | Entity features, geometry, interaction parameters, evolving edges | Correct object segmentation or composition outside training |
| Equivariant geometric operations | Scalar functions and compatible vector/tensor features | Symmetry of a scene with fixed gravity and fixtures |
| Energy, interconnection, dissipation | Energy function, inertia, damping, bounded residuals | True energy, stability with arbitrary residuals, or accurate contact |
| Contact constraints / hybrid modes | Contact geometry, friction, mode probabilities, compliant corrections | A correct rigid-contact approximation for every material |
| Probabilistic rollout interface | State and parameter posteriors, predictive distributions | Calibrated uncertainty or safe decisions |
Proposed design: begin with a small selection of these components, then test each. A single monolithic network containing every physical formalism would obscure which assumption helps. The possible contribution is a useful combination and an experimentally supported reuse boundary; none of the individual ingredients is new.
03 / STATE UNDER PARTIAL OBSERVABILITYThe robot sees evidence, not the state of the world.
A Markov state contains enough information that the future depends on the present state and action, rather than the entire past. A single image is generally not one. Motion can be hidden, an object can be occluded, and friction can remain unknown until the robot acts. A belief is a distribution over possible hidden states given the interaction history.
The proposed state includes robot configuration, object poses and velocities, geometry, persistent identity, articulated structure, support relations, and a contact graph. It may also need mass and inertia estimates, friction, compliance, deformability descriptors, actuator state, and uncertainty. Some variables are slow parameters; others change quickly. Heat, wear, damage, or a shifting load can invalidate an assumption that parameters stay fixed.
The model does not need to reconstruct everything visible. It needs a decision-sufficient physical state: enough information to rank actions for the relevant task family. A box's printed logo may be irrelevant to pushing. The offset of its center of mass may be decisive. The difficulty is that the next task can make a previously irrelevant detail important.
Technical note: a belief, a transition, and an observation model
Let \(o_t\) be a synchronized multimodal observation at discrete time \(t\), \(a_t\) the commanded action over the next interval, \(z_t\) a latent physical state, \(\theta_t\) hidden physical parameters, and \(c_t\) a discrete contact or hybrid mode. Then:
A practical filter approximates this posterior with particles, a mixture distribution, or a recurrent latent representation with probabilistic heads. Given a previous belief, it first propagates through an action-conditioned transition, then updates with a likelihood for the new observation. Denote the combined hidden state by \(s_t=(z_t,\theta_t,c_t)\):
Here \(p_\phi\) is the shared transition with learned parameters \(\phi\), \(p_\psi\) is an observation model with parameters \(\psi\), and \(\theta_e\) describes the embodiment. The integral includes summation over discrete modes; the proportionality constant normalizes the belief. This equation assumes the augmented state is Markov and observations are conditionally independent of the past given that state. Unknown actuator delay must therefore appear in the augmented state or action history.
A learned encoder is not automatically a Bayesian filter, and a latent variance is not automatically physical uncertainty. Without an identifiable observation model, different latent states may explain the same observations. Training must connect the proposed belief to measured future geometry, forces, events, or other decision-relevant targets.
Representation is a tradeoff, not a ladder of sophistication
| Representation | What it makes easier | What it can hide |
|---|---|---|
| Pixels / video | Learning from broad visual data; predicting appearance | Forces, hidden properties, causal action alignment |
| Compact learned latent | Fast rollout and task-relevant compression | Units, constraints, identity, state sufficiency |
| Point cloud / 3D / 4D geometry | Shape, spatial relations, geometry through time | Occluded surfaces, material properties, contact forces |
| Objects / scene graph | Identity, relations, variable object count | Segmentation errors; distributed deformation |
| Explicit physical variables | Units, mechanics, interpretable constraints | Unmodeled effects; expensive or ambiguous identification |
Latent planning is established, while recent surveys organize manipulation models across visual, latent, geometric, and physical representations.[19][14] A plausible starting point is mixed: explicit poses and velocities for rigid objects, learned shape and material descriptors, and a recurrent belief for unobserved variables. The research question is how small that state can be before its compression destroys useful counterfactual predictions.
04 / OBJECTS & RELATIONSReuse the interaction, not the whole scene.
A dynamic interaction graph represents entities as nodes and possible physical relations as edges. A node can be a robot link, rigid object, fixture, or a set of particles representing a deformable body. Edges encode relative pose, proximity, contact mode, attachments, or constraints. An edge can appear when objects approach and disappear when they separate.
Applying the same learned interaction operator to object–table, gripper–object, and object–fixture pairs creates an opportunity for compositional transfer. Graph-based learned simulators demonstrate the usefulness of shared local computations for physical systems.[27] They do not establish that a model trained on a few contacts will generalize to arbitrary clutter or new embodiments.
03 / One scene, several kinds of relation
Node
Pose · velocity · geometry · material descriptors · belief about hidden properties.
Edge
Relative geometry · potential contact · active mode · constraint · interaction features.
A new part reuses operators. Its geometry and parameters still need grounding.
Technical note: local messages do not make contact local
Here \(z_i\) is node \(i\)'s state, \(e_{ij}\) its edge features with neighbor \(j\), \(\mathcal N(i)\) the neighbor set, \(M_\phi\) a shared message function, \(m_i\) aggregated interaction features, \(D_\phi\) a shared dynamics function, and \(u_i\) the input acting on the node. This is a candidate smooth-mode computation, not a complete contact solver. Summation is invariant to neighbor ordering; that alone enforces neither force balance nor momentum conservation.
Simultaneous contacts can couple distant nodes through a rigid chain. A shallow message-passing network may not resolve that coupling. More iterations, multiscale communication, or an explicit constraint solve may be necessary. Contact graph mistakes can dominate all improvements in the dynamics operator. Compare oracle graphs against inferred graphs to expose this failure mode.
05 / STRUCTURED CONTINUOUS DYNAMICSMechanics supplies useful constraints. It does not supply the whole model.
A neural ordinary differential equation learns a continuous-time rule for how state changes.[22] One can restrict that rule using mechanics. Hamiltonian networks derive conservative motion from an energy-like scalar; Lagrangian networks provide an alternative based on positions and velocities.[23][24] Neither is a sufficient description of all manipulation.
Proposed design: use a port-Hamiltonian-style substrate where smooth mechanical state is meaningful. “Port” means a channel through which energy enters or leaves, such as an actuator. Represent energy exchange, damping, and inputs explicitly; reserve residual modules for inadequately modeled effects. Use a separate treatment for nonsmooth contact and a different state or formalism where deformation demands it.
05 / A candidate smooth-mode dynamics block
Energy H
Learned or partly analytic kinetic and potential structure.
Interconnection J
Skew-symmetric coupling; no energy creation by this term.
Dissipation R
Positive-semidefinite loss, within the chosen model.
Actuation Gu
Realized input from the embodiment model.
Residual fres
Learned discrepancy; may invalidate energy guarantees.
Integrator + contact
Advance smooth state; resolve constraints and mode changes.
State and inferred parameters → structured vector field → numerical update → next state.
Technical note: Hamiltonian structure, passivity, and the residual problem
For a smooth conservative mechanical system in canonical coordinates, let \(q\) be generalized position, \(p\) conjugate momentum, \(x=(q,p)\), \(T\) kinetic energy, and \(V\) potential energy:
This assumes a valid canonical state and an autonomous, differentiable Hamiltonian. Constant energy does not imply a closed orbit, a correct trajectory, or stability. A learned \(H\) need not be the true physical energy, especially when its coordinates are latent.
Here \(J=-J^\top\) is skew-symmetric interconnection; \(R=R^\top\succeq0\) is dissipation; \(\nabla H\) is the state gradient; \(G\) maps realized actuator input \(u\) into state dynamics; and \(f_{\mathrm{res}}\) is a residual that may depend on additional latent context \(z\). These quantities can also depend on inferred physical parameters and mode, suppressed for readability. For a time-independent \(H\), the chain rule gives:
The output \(y\) is conjugate to \(u\), so \(y^\top u\) is input power in a physically grounded realization. Skew-symmetry removes the \(J\) contribution. With no residual, the smooth dynamics cannot increase \(H\) except through the input port. That is an energy-balance property, not an unconditional stability theorem. With an arbitrary residual, even that statement no longer holds.
One option is to constrain residual power or place uncertainty around it; another is to accept loss of passivity and measure the consequences. Positive-semidefinite parameterization of \(R\), for example through a matrix factor, is architectural enforcement. A penalty encouraging it is only soft regularization. Nonnegative storage, appropriate equilibrium structure, and further conditions are needed for stability claims. Existing stable port-Hamiltonian work proves guarantees for its specified architectures; those guarantees do not automatically extend to this hybrid, partially observed proposal.[26]
A skew matrix alone also does not make an arbitrary interconnection a Poisson structure: additional identities are needed if that geometric claim is made. Rotations live on a manifold, not unconstrained Euclidean coordinates. Lie-group models preserve the appropriate configuration structure; Lagrangian descriptions may be more natural when velocities are observed and canonical momenta are not.[25] Explicit time dependence adds another term to the energy balance, and impacts require a separate jump balance. Updating inferred parameters can also change the model's energy estimate; that bookkeeping change must not be mistaken for physical work.
This is where an attractive physical prior can become a liability. A rigid state cannot explain a buckling package. A passive model cannot explain an omitted powered mechanism. If residuals learn almost everything, the energy scaffold may add cost without useful restriction. The relevant comparison is with a capacity- and data-matched model, not with an intentionally weak predictor.
06 / HYBRID CONTACTThe difficult physics happens when the rules change.
A hybrid dynamical system combines continuous evolution with discrete changes of mode. A part moves freely, hits a fixture, slides along it, sticks, then loses support. A grasp closing is an actuator event; a stable grasp additionally requires appropriate contact and force. An insertion can move from unconstrained motion to several simultaneous constraints.
04 / Contact is a branching process
Guard: when does a transition occur?
Reset: what changes across that event?
Mode dynamics: what happens between events?
The challenge is not merely predicting that two shapes touch. Contact timing depends on small geometric differences. Sticking and slipping can be visually ambiguous. Rare impacts receive little training coverage. Multiple contacts transmit forces through an assembly, while stiff or deformable interactions couple fast local changes to slower motion. A predicted average between “slips” and “sticks” may correspond to neither physically possible outcome.
Technical note: guards, resets, complementarity, and friction
Mode \(c\) selects the within-mode dynamics \(f_c\). A guard function \(g_{c\to c'}\), together with its crossing direction and feasibility conditions, signals transition to mode \(c'\). The reset map \(\Delta\) takes the state immediately before the event, \(x^-\), to immediately after, \(x^+\). At a rigid impact, positions are usually continuous while momentum jumps. A guard can also depend on input, time, or hidden parameters; the compact notation omits these possibilities.
For an ideal rigid, nonadhesive unilateral contact, \(\varphi(q)\) is signed gap and \(\lambda_n\) the compressive normal contact force. Bodies cannot overlap; the surface cannot pull them toward itself; positive force requires zero gap. Zero gap does not require positive force. These position-level conditions are necessary, not a complete dynamics law: velocities, accelerations, other constraints, and an impact law determine the force or impulse.
Here \(\lambda_t\) is tangential contact force and \(\mu\geq0\) the Coulomb friction coefficient. The cone bounds admissible force. Sliding additionally needs a direction opposing relative tangential motion, while sticking requires compatible zero slip. At impacts, impulse variables and a restitution or dissipation rule replace an ordinary finite-force update. Coulomb friction neglects many effects, including adhesion, hysteresis, and velocity-dependent materials.
Rigid-contact complementarity can be nonunique or ill-conditioned. Time-stepping schemes provide an alternative to resolving every impact as a separate event.[29] Smooth compliant contact eases some optimization problems but introduces stiffness and approximation error. Differentiating through contact may require implicit differentiation, event sensitivity, or a surrogate; gradients can be unreliable exactly at mode boundaries. Choose the approximation for the task and report constraint violations.
Recent work changes the baseline
ContactWorld studies representations for vision-tactile predictive planning; its revised 2026 preprint emphasizes spatial structure, temporal continuity, and compatible modalities. Its lesson is that adding a tactile channel alone is insufficient.[16] Dream-Tac jointly models actions, future visual observations, and tactile dynamics, using contact-aware fusion.[18] Both belong in the comparison set for contact prediction, rather than being treated as peripheral perception work.
WorldContact, a September 2026 preprint, addresses deformable-object interaction and uses its learned dynamics to expand policy-training data.[17] This demonstrates a relevant architectural route: a predictive model can help a VLA during training without becoming its runtime planner. These papers support investigating contact-aware representations; they do not validate the complete Physical Firmware factorization.
07 / THE EMBODIMENT ADAPTERA small identification problem—if the hypothesis holds.
System identification infers a system's dynamics or parameters from input–output observations. The adapter should perform identification as well as encoding. It connects coordinate frames, sensor calibration, morphology, action interfaces, and the physical model. A position target, velocity command, and torque command are not interchangeable interventions.
The shared model has learned parameters φ. The robot has embodiment variables θe: kinematic structure, actuator gains and lag, joint friction, compliance, gripper geometry, sensor transforms, control frequency, and small residual modules. A morphology description can supply known structure; controlled interactions infer uncertain quantities. The desired result is to update θe while initially holding φ fixed.
Amortized identification trains an inference model across many systems so a new interaction history can produce a parameter estimate without fitting everything from scratch. Rapid Motor Adaptation is a relevant example of history-based online adaptation in locomotion; it is not evidence that a manipulation predictor will transfer across arbitrary robots.[31]
Calibration should ask identifiable questions
Free-space motion at different speeds can reveal lag and actuator response. Gentle surface contact can ground geometry and compliance. Pushing, lifting known loads, opening and closing the gripper, controlled slip, and bounded force application reveal different combinations of parameters. Calibration must obey the robot's validated limits and use force-limited procedures appropriate to its hardware.
Not every parameter is separately identifiable. A slow push may reveal a combined friction effect but say little about inertia; actuator tracking error can masquerade as an object-dynamics error. Keep parameter uncertainty when multiple explanations remain consistent. Choose additional probing actions for information only when their expected benefit justifies their risk and time.
Technical note: commands are not physical inputs
Here \(a_t\) is a command, \(h_t\) is actuator/controller history including delayed commands, \(A_{\theta_e}\) updates that history, and \(B_{\theta_e}\) predicts the physical input \(u_t\) realized over the next interval. The dynamics core consumes that realized input or its distribution. A memoryless map is appropriate only when lag and controller state can be neglected at the modeling time scale.
A posterior over \(\theta_e\) can be inferred from calibration observations and carried into rollout. A robot-specific residual should be capacity-limited and audited. If the adapter learns the entire task or absorbs the whole world model, calling the remaining core “shared” does not demonstrate useful transfer. Report adapter capacity, data, compute, and retained performance on previous embodiments.
Can a new embodiment be characterized with a small amount of structured interaction, at lower total cost than adapting a strong generalist policy or retraining the predictor?
08 / UNCERTAINTYA distribution over futures, not a confidence badge.
A robot can be uncertain because it cannot see the contact, because the friction is unknown, or because its model has never encountered this material. Those are different reasons to hesitate. They may call for a better view, a probing interaction, a shorter planning horizon, a conservative action, or a handoff.
| Uncertainty | Meaning | Useful response |
|---|---|---|
| Aleatoric | Stochasticity or ambiguity remaining under the chosen sensing and state representation | Propagate possible outcomes; improve sensing when hidden state is recoverable |
| Epistemic | Lack of knowledge in the learned model | Detect unfamiliar conditions; acquire informative data |
| Mode | Whether an impact, stick, slip, or release will occur | Keep distinct branches instead of averaging them |
| Parameter | Unknown mass, friction, compliance, or actuator response | Maintain and update a parameter belief |
| Horizon | How predictive reliability changes with look-ahead time | Evaluate calibration separately at each useful horizon |
These are overlapping axes, not mutually exclusive error bins. Hidden friction creates parameter uncertainty and may change the contact mode. “Irreducible” aleatoric uncertainty is relative to the available information: adding tactile sensing can make previously hidden variation predictable.
06 / One action, several plausible consequences
Probabilistic ensembles with trajectory sampling, as in PETS, offer a practical starting point.[20] An ensemble represents disagreement across fitted models; stochastic heads represent within-model variability. Neither guarantees that all models will disagree on an unfamiliar case. Out-of-distribution (OOD) detection asks whether current inputs or transitions differ from the training conditions, not whether a prediction is certainly wrong.
Recent horizon-calibrated world-model work explicitly trains temporal uncertainty. Its action-free video pretraining setting is relevant but different from torque-conditioned contact prediction.[34] Physical Firmware would need held-out coverage tests at multiple horizons, under contact transitions and under the planner's own selected actions. A variance forced to increase is not evidence of calibration, and increasing variance monotonically is not a universal physical law.
Technical note: the rollout distribution and calibration contract
Here \(K\) is the number of future transitions, \(b_t\) the current belief, \(\theta\) uncertain environment parameters assumed constant within this short rollout, and \(\theta_e\) embodiment parameters. The sequence has \(K\) actions, not \(K+1\). If environment parameters vary, predict their sequence as well. Integrate over uncertain embodiment parameters rather than treating their estimate as exact.
Rollouts should retain temporal correlations. Sample one model and a persistent parameter hypothesis for a trajectory where appropriate; independently resampling mass every step would represent a different physical process. Integrate mode probabilities with state uncertainty rather than attaching an unrelated confidence scalar after prediction.
Calibration means stated uncertainty matches empirical frequencies on a specified distribution. Check interval or region coverage together with sharpness: a region containing everything has coverage but little decision value. Conformal methods can calibrate sets under assumptions such as exchangeability; correlated robot trajectories and adaptive action selection violate simple exchangeability. Extensions exist, but arbitrary deployment shift is not covered for free.[33]
Optimization searches for low predicted cost. It can therefore select an action because the model is wrong there. Evaluate disagreement and error on optimized candidate actions, not only random held-out trajectories. Rejecting uncertain plans is useful only if uncertainty actually predicts their failures.
09 / TRAINING PROGRAMLearn broad structure. Then pay close attention to reality.
A plausible program combines simulation, synchronized real interaction, and deliberate embodiment variation. Simulation provides privileged state labels; real data exposes model mismatch. Cross-embodiment training tests whether robot identity has become an accidental shortcut. These stages can overlap; they describe responsibilities, not a proven recipe.
08 / Proposed training pipeline
- Structured simulation pretraining
Varied bodies, contacts, materials, sensors, and action realization. Learn predictive structure with privileged labels.
- Multimodal real-world prediction
Synchronize images, geometry, proprioception, force, touch, commands, and realized actuator state. Learn the discrepancy.
- Cross-embodiment training
Share core operators; vary robot morphology and interfaces. Hold out whole embodiments.
- Deployment identification
Freeze the core initially. Infer embodiment parameters; fit limited calibration and residual modules.
- Continual learning with regression gates
Curate deployment residuals; replay prior tasks; verify constraints and retained performance before updates.
A / Simulation teaches a family of possible interactions
Randomize mass, inertia, friction, compliance, damping, geometry, object count, morphology, sensor noise, and action latency. Use rigid-body simulation first; add deformable or differentiable simulation when the task requires it. Ground-truth pose, force, mode, and parameter labels help separate failures of perception from failures of dynamics.
Differentiable simulators can support parameter fitting and gradient-based planning; DiffTaichi is one established route to differentiable physical computation.[32] Differentiability does not remove the reality gap. Simulation pretraining should teach variation and structure, not certify that contact labels, friction, or sensor models match hardware.
B / Real prediction needs synchronized interventions
Record RGB, depth, joint positions and velocities, measured or estimated torques, wrist force/torque, tactile signals where available, commanded actions, actuator state, timestamps, outcomes, and failures. Record units, frames, calibration versions, sampling intervals, and missing-modality masks. Command timestamps and observed execution timestamps are different data.
Predict future features, geometry, motion, contact events, and forces. Not all real data has all labels: mask unavailable terms, use sensor likelihoods, and distinguish estimated contact annotations from instrumented ground truth. Resample with care around impacts; interpolating a discontinuity can erase the event that matters.
Passive videos supply visual and temporal priors but generally do not identify the effect of a robot command. Observed actions can be correlated with hidden causes. Controlled interventions, varied action coverage, and explicit action realization are needed before interpreting a predictor causally. A model that imitates the dataset's typical future is not necessarily accurate for a planner's counterfactual action.
C / Reuse must be exercised during training
Train with several morphologies, grippers, camera configurations, and control interfaces where feasible. Normalize units and frames without erasing meaningful differences. Test held-out combinations and a wholly held-out robot, not merely a new episode from a familiar robot. Public cross-embodiment datasets are valuable, but their sensor and force coverage must be audited for this objective.[7]
D–E / Adaptation and continual learning are separate budgets
At deployment, first update embodiment parameters, small residuals, and calibration layers. If that fails, allow a separately reported core-update condition. Continual learning can later incorporate new residuals and failure cases using replay, constrained parameterizations, and a fixed regression suite. Catastrophic forgetting—losing previous capabilities while learning new ones—must be measured across old robots and tasks. A held-out benchmark is not a replay buffer.
10 / TRAINING OBJECTIVESOptimize the predictions the controller will actually use.
A model can predict the next frame well and fail after repeated rollout. Training should expose it to its own predicted states, to contact transitions, and to several horizons. Latent overshooting compares multi-step predicted latent distributions with later inferred states; it was used in PlaNet.[19] Scheduled rollout training gradually replaces ground-truth context with predicted context. Neither automatically solves the difference between training data and a planner's chosen actions.
Technical note: one candidate multi-objective loss
Each nonnegative \(\lambda\) weights a term after units and scales are normalized. The terms below are a design menu, not nine objectives known to be jointly optimal.
| Term | Purpose and limitation |
|---|---|
| Rollout, Lroll | Multi-horizon likelihood or proper predictive score for future observable targets; not just next-state error. |
| Contact, Lc | Mode/event likelihood and timing; handle imbalance without destroying probability calibration. |
| Force, Lf | Force/torque or impulse prediction with sensor noise and bandwidth modeled. |
| Geometry, Lg | Pose, distance, shape, identity, and geometric consistency for available labels. |
| Equivariance, Le | Consistency under valid transformations when symmetry is not enforced exactly. |
| Structure, Ls | Penalty for unenforced constraints, inappropriate energy gain, or invalid contact. Avoid duplicate constraints. |
| Uncertainty, Lu | Proper distributional scores for additional probabilistic heads; omit duplicate NLL already in rollout loss. |
| Calibration, Lcal | Validation-tuned calibration or coverage surrogate, with a separate untouched evaluation set. |
| Representation, Lr | Temporal, cross-modal, or latent consistency without collapsed representations. |
Here \(\mathcal D\) is the training trajectory distribution, \(K\) the maximum rollout horizon, \(w_k\geq0\) horizon weights, and \(y\) an observed target such as pose, geometry, or sensor features. This is a sum of marginal predictive negative log likelihoods, not the joint likelihood of the entire trajectory. Missing labels require masks. The distribution must arise from the same rollout and observation interface used during planning.
A learned latent target needs an anchored decoder or a noncollapse objective; reducing latent error by shrinking all features is not learning physics. A likelihood must specify noise and units. If the predictor only produces samples, use an appropriate sample-based proper score instead of claiming a tractable NLL. A hard architectural constraint needs no penalty to “enforce” the same property. Measure whether each remaining loss improves downstream decisions.
Use rollout corruption and diverse action sequences to expose compounding error. Collect new interactions where the current planner fails, with a separate safety and data-governance process. Evaluate on independently collected episodes and splits by object, parameter combination, and embodiment. Randomly splitting adjacent frames would make the apparent generalization nearly meaningless.
11 / SYMMETRYTransform the problem consistently.
Equivariance means that transforming an input transforms the output in the corresponding way. A vector force should rotate when the coordinate frame rotates; a scalar mass should not. Equivariant graph networks provide tools for building such behavior into learned computations.[28]
The world is not indiscriminately SE(3)-symmetric. SE(3) is the group of three-dimensional rotations and translations. Rotating a scene while keeping gravity fixed can change what happens. A fixed robot base, asymmetric gripper, fixture, camera, and external field also matter. A change of coordinates transforms those quantities too; physically rotating only the object is a different intervention.
Technical note: gravity-conditioned equivariance
Here \(F\) predicts a physical state or compatible vector field, \(g\) is a valid rigid transformation, and \(\gamma\) collects gravity and other environmental context. The dot denotes the appropriate group action on each quantity: vectors rotate, points rotate and translate, and invariant scalars stay fixed. Actions must transform according to their meaning; joint commands do not transform like Cartesian forces.
For a physical transformation that keeps \(\gamma\) fixed, only transformations preserving that context are symmetries. Observation encoders have additional visibility and camera constraints. Apply equivariance to selected geometric components and test with transformed gravity, robot, and fixture descriptions where required. A generic SE(3) augmentation of images is not the same guarantee.
The expected benefit is fewer examples needed to learn the same relation in different frames. The risk is forbidding a real asymmetry. Compare architectural equivariance with valid data augmentation and with an equally expressive unconstrained model.
12 / NUMERICAL INTEGRATIONThe integrator is part of the model.
A continuous-time dynamics equation is not a trajectory. A numerical integrator advances it in finite steps. Explicit Euler is simple but can accumulate severe error. Runge–Kutta methods improve local accuracy for smooth dynamics. Symplectic and variational methods preserve selected geometric structure; for appropriate smooth Hamiltonian problems, their long-time behavior can be much better than a generic discretization.[30]
They do not universally conserve exact energy, resolve stiff contact, or guarantee accurate control. Near an impact, event detection and reset handling may matter more than smooth-step order. Persistent constraints can require stabilization or projection. Stiff compliant contact may need implicit updates or small time steps. Adaptive stepping changes computational cost and can interfere with geometric guarantees unless designed for them.
Proposed design: train and evaluate the actual discrete rollout used at deployment, including its contact solver, tolerances, and time step. Compare equal wall-clock budgets as well as equal step counts. A model trained with one solver may compensate for that solver's error and fail when integrated differently. Report both physical constraint violations and task outcomes.
13 / PHYSICAL FIRMWARE + VLAA consequence model can complement a generalist policy.
The useful shorthand is: VLAs learn what action to take; world models predict what may happen; simulators encode an explicit approximation of what should happen. Physical Firmware asks whether selected architectural physics improves a reusable learned predictor. These are roles, not mutually exclusive model species.
Flow and diffusion policies learn distributions of action sequences, which is valuable when several behaviors can satisfy the same instruction. Diffusion Policy and π0 establish two important action-generation lineages.[8][1] A distribution over good actions is different from a distribution over the consequences of an arbitrary proposed action. A single architecture can learn both.
07 / Five possible integration patterns
- A / Direct policy
Observation + instruction → VLA → actions
The strong baseline: reactive or history-conditioned behavior, without an exposed rollout service. - B / Consequence critic
VLA proposals → firmware predictions → select / modify → controller
Compare candidate consequences. Rejection is only as good as model calibration and coverage. - C / Semantic goal + MPC
VLA subgoal → cost / constraints → firmware + action search → controller
A model-predictive controller searches locally; translating language into a valid objective is still a separate problem. - D / Training infrastructure
Firmware → counterfactuals / hard cases → policy training → fast VLA
Use predicted experience offline, with real validation and controls on model bias. - E / Hybrid runtime
VLA sequencing ⇄ local physical planner ⇄ stabilizing controller
Separate time scales; use measured feedback to correct all three levels.
Predictive auxiliary objectives already appear in policy learning. GR00T N1.5 adds Future LAtent Representation Alignment (FLARE) to action learning, aligning representations with future embeddings rather than generating future frames.[10] π0.7 uses richer conditioning, including visual subgoals produced by a lightweight world model.[4] These are concrete evidence that prediction and action learning are converging.
The distinction to test is therefore sharper than “has a world model.” Does an explicit, reusable physical rollout interface—with identified action realization, contact structure, and calibrated uncertainty—add value beyond the predictive representations already inside a policy? The answer could be no. An auxiliary loss might capture the useful structure at lower runtime cost.
14 / PLANNINGPredict. Choose. Act briefly. Observe again.
Model-predictive control (MPC) repeatedly evaluates action sequences over a finite horizon, executes the first action or short prefix, updates its state estimate, and replans. A sampled optimizer can compare future outcomes without differentiating through every contact. Gradient-based planning is another option when the model and its derivatives are reliable.
Cost can include distance to a goal, time, energy, force, constraint violation, and task failure. Risk-sensitive planning accounts for the distribution of outcomes rather than just its mean. A rare drop or damaging contact may deserve more weight than a small improvement in average speed. Uncertainty penalties should be validated for usefulness, not selected because cautious trajectories look reassuring.
Technical note: risk-aware MPC and its limits
The candidate sequence \(\mathbf a=(a_t,\ldots,a_{t+K-1})\) belongs to \(\mathcal A_K\), the set respecting command and actuator limits. The random trajectory \(\tau\) comes from the rollout distribution; \(C\) is accumulated task cost; \(U\) is a chosen uncertainty penalty; and \(\beta\geq0\) sets its weight in compatible units. This objective is a proposed design, not a theorem that uncertainty penalties produce safety.
A model-based chance constraint can bound predicted probability of leaving a chosen admissible set \(\mathcal S\), at tolerance \(\varepsilon\). The real-world bound is only as valid as the model and calibration assumptions. Tail-cost objectives such as conditional value at risk emphasize bad outcomes; they also require enough samples of those outcomes to be meaningful. Independently enforced robot limits and an explicit fallback remain necessary.
The open-loop candidate sequence does not include every possible future observation and recovery action. Replanning partly compensates, but full belief-space or dual-control planning—which chooses actions both to control and to learn—costs more. Do not count the benefit of future feedback twice in an open-loop rollout.
Latency is an architectural constraint
High-frequency motor stabilization belongs in a suitable controller. Medium-horizon physical planning can run more slowly; semantic sequencing can run more slowly still. Their exact frequencies depend on the robot, task, sensors, and communication path. A predictor taking seconds cannot sit synchronously inside a loop requiring a new decision every 100 milliseconds; this is a timing example, not a measured system specification.
Measure end-to-end deadlines, including state estimation, candidate generation, integration, uncertainty sampling, and transport—not just one neural forward pass. Warm-start planning, batch candidates, shorten horizons, distill policies, and use hierarchical models when useful. If a deadline is missed, use a defined safe hold or fallback appropriate to the hardware. Model-based RL also teaches that short, trusted rollouts may outperform extensive use of a biased model.[21]
Proposed interface: the data contract a planner would need
# Interface sketch — no implemented library is claimed.
belief = estimator.update(history, robot_description)
grounding = adapter.infer(calibration_history, belief)
prediction = core.rollout(
belief=belief,
command_sequences=candidates,
embodiment=grounding,
sample_times=planning_times,
deadline=remaining_planning_budget,
)
# Return correlated trajectory samples, contact modes,
# forces, parameter hypotheses, validity diagnostics,
# reference frames, units, timestamps, and model version.
plan = planner.evaluate(prediction, goal, constraints)
controller.execute_first_prefix(plan)
# Compare actual consequences; update belief and replan.A prediction response should specify whether its uncertainty is calibrated on comparable conditions, which assumptions failed, and whether computation met its deadline. An invalid prediction is a possible output, not an exception to hide.
15 / WORLD MODEL VS. SIMULATORThe useful distinction is what each model commits to.
A classical simulator specifies geometry, equations, parameters, and a numerical approximation. Its state is inspectable and interventions explicit. Its weaknesses are imperfect identification, simplified contact and materials, and the simulation-to-reality gap. It can already incorporate learned parameters or residuals.
A generative video world model predicts rich visual futures. Broad visual training can cover scenes and behaviors difficult to author manually. But visually plausible motion is not necessarily accurate under a particular actuator command, and a convincing image does not establish correct force or contact. Action-conditioned variants need a precise account of what the action channel means.
A latent dynamics model predicts compact hidden representations. It can be efficient for planning without rendering images, but its state may be hard to inspect or constrain. A differentiable simulator provides derivatives through its approximate dynamics, enabling identification or optimization; differentiation is a computational capability, not a guarantee of realism.
Physical Firmware belongs in this overlapping space. It proposes learned dynamics plus selected physical structure, real-world residuals, explicit uncertainty, and a reusable embodiment interface. If implemented as a structured learned simulator, that description is entirely appropriate. It should earn its separate name through useful transfer, not through a claim to a previously empty category.
- Physical consistency ≠ prediction accuracy.
- Prediction accuracy ≠ planning usefulness.
- Visual realism ≠ physical accuracy.
- One-step accuracy ≠ stable rollout ≠ successful control.
- Simulation transfer ≠ real-world transfer.
- Architectural prior ≠ scientific truth.
16 / WHAT IS ACTUALLY UNIVERSAL?Universal laws do not imply a universal learned representation.
A rigid-body state can transfer across many objects while failing on cloth. A contact operator useful for a parallel gripper may miss the distributed compliance of a hand. Tabletop manipulation and locomotion share mechanics but differ in reachable states, contact topology, sensing, and the consequences of failure.
The transferable unit might be mechanics, a set of contact primitives, a geometric representation, a parameter-inference procedure, or simply broad pretraining. Those alternatives make different predictions. If a pretrained unstructured model adapts equally well, useful reuse exists, but the claim for explicit architectural physics weakens.
Speculation: the eventual design may be a hierarchy or a learned mixture of specialized models: rigid-body, articulated-body, deformable-object, fluid-interaction, and tool-interaction modules. Such a mixture needs a routing rule, compatible state and force interfaces, and uncertainty about choosing the wrong model. Joining individually plausible modules does not guarantee that their combined energy and contact behavior is consistent.
Start with a declared domain. “Reusable within contact-rich rigid tabletop manipulation” is a meaningful claim. “Universal physical intelligence” is not a justified extrapolation from it. Cross-object transfer, cross-task transfer, and cross-embodiment transfer must be reported separately.
17 / THE SEPTEMBER 2026 LANDSCAPEThe categories are converging.
Physical Intelligence's progression matters because it strengthens the behavioral baseline. π0 combines broad robot data with a flow-based action model; π0.5 addresses open-world generalization; π*0.6 uses learning from experience and reinforcement-learning post-training; π0.7 adds steerability and reports compositional and cross-embodiment behavior.[1][2][3][4] These results concern the authors' evaluated settings; they are not a claim of arbitrary task competence.
Its 2026 Multi-Scale Embodied Memory work adds long- and short-term history, while RL Token work uses compact VLA representations for online policy improvement.[5][6] Memory matters because observation history can disambiguate state. Online adaptation matters because frozen imitation is no longer the only policy baseline. Neither, on its own, implies an explicit posterior over physical parameters.
NVIDIA's GR00T N1 is a generalist humanoid policy; N1.5 couples policy training to future representation alignment. Later N1.6 and N1.7 releases extend the available model family; the official repository identifies N1.7 as its current release and describes relative end-effector actions shared across human and robot data.[9][10] Action representation itself can support transfer. The other important precedent is joint prediction and action learning. An auxiliary embedding objective should not be mistaken for a fully identified, externally queryable contact simulator.
NVIDIA Cosmos develops world foundation models for physical AI. The original platform emphasizes large-scale world modeling; Cosmos 3's 2026 technical report describes an omnimodal model spanning understanding, generation, action, and forward/inverse dynamics.[11] This makes a strict “video model versus policy” taxonomy increasingly inadequate. Capabilities described in a technical report and released model availability should still be distinguished.
Google DeepMind's Gemini Robotics work joins embodied reasoning to action, while Gemini Robotics ER 2, announced in July 2026, emphasizes video-based progress monitoring, task orchestration, and delegation to lower-level execution systems.[12] Genie 3 demonstrates interactive generated environments; such environments are relevant to experience generation, but visual interactivity alone does not establish calibrated robot-contact dynamics.[13]
Two 2026 surveys help organize this convergence: Wang and colleagues distinguish representations and how prediction is connected to action; Kirchner, Purschke, and Knoll foreground uncertainty and closed-loop control.[14][15] The comparison below is an architectural map, not a benchmark ranking. Entries marked “varies” depend on implementation and training data.
| Approach / examples | Output | Physical structure | Action conditioning | Cross-embodiment | Contact | Uncertainty | Primary role |
|---|---|---|---|---|---|---|---|
| VLA / π family; Gemini Robotics | Actions / action sequences | Mostly learned; architecture varies | Generates actions from context | Trained / evaluated in multiple settings | Often implicit; sensor-dependent | Action diversity is not calibrated outcome uncertainty | Generalist behavior |
| VLA + predictive objective / GR00T N1.5 | Actions + future representation supervision | Learned visual/action representations | Joint training; exposed rollout API not implied | Multi-embodiment data | Implicit in described objective | Not certified by auxiliary prediction | Representation and policy learning |
| Generative world foundation model / Cosmos; Genie 3 | Video / multimodal futures; some models also actions | Learned; geometric conditioning varies | Varies: control, video, or robot-action channels | Requires action and embodiment grounding | Visual plausibility is insufficient | Sample diversity; calibration must be tested | Generation, predictive infrastructure |
| Classical physics simulator | Physical trajectories | Explicit equations and geometry | Explicit controls / forces | New robot models and parameters | Chosen numerical contact law | Usually added through parameter/noise models | Simulation, control, synthetic data |
| Differentiable simulator / DiffTaichi | Trajectories and derivatives | Explicit or hybrid | Explicit, model-dependent | Re-identification / remodeling | Gradients depend on formulation | Not automatic | Identification, optimization |
| Learned latent dynamics / PlaNet; PETS | Latent or state futures | Learned transitions; priors vary | Explicit candidate actions | Not automatic | Often learned implicitly | Stochastic state / ensembles | Planning and model-based RL |
| Physics-structured dynamics / HNN; port-Hamiltonian models | State derivatives / rollouts | Selected energy / geometry constraints | Depends on model; ports in controlled variants | Must be demonstrated | Separate extension usually needed | Depends on implementation | Dynamics learning, control |
| Physical Firmware / proposal | Grounded probabilistic physical futures | Selected graph, geometry, dynamics, contact priors | Commands through identified realization | Central hypothesis | Explicit modes / constraints + learned corrections | Belief, parameters, modes, model uncertainty | Reusable prediction layer |
On narrow screens, scroll the comparison horizontally. “Not automatic” is not a claim that a capability is impossible or absent from every system in that category.
18 / THE DECISIVE EXPERIMENTMeasure the cost of reaching the same reliability.
The central prediction is a leftward shift: less new real-world experience to reach the same reliability after a physical change.
Start with contact-rich tabletop manipulation: pushing, repositioning, constrained sliding, stacking, tray placement, pick/place under mass variation, and later peg or connector insertion. Hold task semantics approximately constant while changing physical conditions. A better nominal score is useful, but it does not by itself demonstrate reusable physical knowledge.
09 / The claim to test — conceptual curves, not results
Baselines that could actually disprove the claim
| Arm | System | What the comparison asks |
|---|---|---|
| A | Task policy trained from scratch | What is the total gain over task-only learning? |
| B | Pretrained VLA / behavioral policy, fine-tuned; include a viable online-RL variant | Does it beat strong behavioral reuse and adaptation? |
| C | Unstructured learned world model + the same planner | Does architectural structure add anything beyond predictive pretraining? |
| D | Simulator-trained policy; also a calibrated simulator + MPC where feasible | Is learning the proposed core better than existing simulation and identification? |
| E | Established physics-informed / structured dynamics model | Does the proposed combination add value over known structure? |
| F | Physical Firmware candidate + MPC | Does its prediction layer reduce adaptation effort? |
| G | Physical Firmware representation + learned task policy | Does reuse help without expensive runtime search? |
| H | Ablations: no equivariance, contact structure, energy structure, uncertainty, or shared core | Which component causes any improvement? |
Hold observation quality, action interface, demonstration quality, task data, simulator access, planner budget, and evaluation episodes constant wherever meaningful. Match capacity and training compute for controlled model comparisons. A large pretrained VLA cannot always be matched exactly; report both a controlled scientific comparison and a practical best-available-system comparison, including inherited data and compute.
Plot reliability against real interaction time and episode count at several adaptation budgets. Record all failed attempts and calibration, not only demonstrations used for gradient updates. Report upstream pretraining separately and show amortization across deployments; an expensive reusable core has not saved data in total merely because downstream fine-tuning is small.
Pre-register the reliability target, minimum practically meaningful saving, success tolerance, action deadline, exclusions, and stopping rule. Use paired physical conditions, multiple training seeds, and uncertainty intervals clustered by object or session. Choose the number of trials with a power analysis based on pilot variability. If a method never reaches the target, report that rather than extrapolating its curve.
19 / DISTRIBUTION-SHIFT TESTSChange what the robot sees separately from what the world does.
A distribution shift is a change between training and deployment conditions. A new camera angle tests perception. A hidden mass change tests dynamics inference. Combining them immediately makes it difficult to identify the reason for success or failure.
Perceptual shift
Camera pose, lighting, background, visual texture, depth noise, occlusion, and sensor dropout. Hold physical parameters fixed when possible.
Physical shift
Mass, friction, geometry, object size, center of mass, compliance, and support configuration. Keep visual appearance controlled where possible.
Compositional shift
Object count, clutter, starting pose, combinations of familiar objects, contact topology, and changed fixtures. Some shifts affect both perception and dynamics.
Embodiment / interface shift
New arm, gripper, controller, latency, control frequency, or sensor suite. Report these separately from changing an object on the same robot.
Distinguish interpolation within trained parameter ranges, held-out combinations, and extrapolation beyond those ranges. Use one-factor tests to diagnose, then joint shifts to test practical robustness. Oracle state and oracle parameter conditions reveal how much failure comes from estimation rather than dynamics. Finally test the full sensor-to-action system: privileged-state results alone do not establish deployment transfer.
20 / METRICSPrediction, control, adaptation, and operating cost.
Success rate is necessary but insufficient. A model might reach the goal by using damaging forces, excessive resets, or slow planning. Conversely, lower prediction error on irrelevant pixels may improve none of the outcomes that matter.
| Level | Report | Interpretation guardrail |
|---|---|---|
| Predictive | One- and multi-step pose/velocity error; rollout NLL where defined; force error; contact event precision, recall, F1 and timing; geometry consistency; constraint violation; applicable energy/dissipation residuals | Disaggregate by contact mode and horizon. Define units, event matching windows, and reference sensors. |
| Uncertainty | Coverage and sharpness; reliability curves; NLL; Brier score for event probabilities; OOD response; calibration by horizon and mode | Score optimized actions too. Broad intervals and large variance are not sufficient evidence of useful calibration. |
| Control | Success, recovery, unsafe failure and intervention rates; force-limit violations; model exploitation; latency distribution; deadline misses; action frequency | Separate model-predicted safety from actual outcomes. Average latency hides missed deadlines. |
| Adaptation | Demonstrations; all interaction time; calibration episodes; wall-clock time; gradient updates; fraction and number of parameters changed; retention on prior robots | Do not classify calibration as “free” or cross-object transfer as cross-embodiment transfer. |
| Operational | Engineering and operator hours; deployment time per new variant; downtime; reset labor; compute cost; hardware/failure cost | A data saving can be outweighed by integration or maintenance overhead. |
Model exploitation should have a concrete diagnostic: compare predicted and realized cost for actions selected by progressively stronger search, within the same bounded action space. If more optimization improves predicted cost while worsening measured outcomes, the planner is finding model defects. Log those trajectories as failures of the prediction–planning system.
Contact F1 and Brier scores answer different questions: event detection quality and probability quality. NLL is meaningful only for a declared predictive density. Energy consistency applies only to modeled components with a meaningful energy balance. None should be used as a universal substitute for control evaluation.
21 / FALSIFICATIONWhat would prove the thesis wrong?
A broad philosophical claim about all possible architectures cannot be refuted by one failed implementation. A scoped engineering hypothesis can. For the declared task family, data budget, and model class, specify an effect large enough to justify the added structure, then give the competing systems a fair chance to match it.
- A matched unstructured world model transfers equally well or better.
- The same reliability needs approximately the same amount of task-specific real interaction once all adaptation is counted.
- Adapter calibration costs as much as policy adaptation or predictor retraining.
- Physical priors systematically exclude real behavior, and residual capacity removes any useful distinction from an unstructured model.
- Uncertainty remains miscalibrated where the planner needs it, or conservative penalties erase the claimed efficiency gain.
- Contact prediction fails in held-out impact, slip, or multi-contact regimes.
- The planner repeatedly exploits model errors despite reasonable uncertainty and rollout controls.
- Cross-embodiment retention is negligible after a fair grounding procedure.
- Long-horizon stability improves but successful control and adaptation cost do not.
- VLA pretraining or predictive auxiliary objectives already obtain equivalent reusable physical knowledge at lower total cost.
An underpowered null result is inconclusive. A repeatable equivalence result with uncertainty intervals narrow enough to exclude the predeclared useful improvement is damaging evidence. If only one component helps—say, contact geometry—retain that result and discard the unsupported claim for the larger architecture. If the best solution is a conventional identified simulator, use it.
Conversely, one successful pushing experiment would support a limited candidate, not universality. The stronger claim needs transfer across physical changes and embodiments, reliable uncertainty under planning, and a measurable reduction in total adaptation effort.
22 / RESEARCH ROADMAPMake each stage earn the next one.
10 / Proposed program, with decision gates
- P0Mathematical prototype
Pendulum, cart pole, spring–mass, articulated chain. Verify parameterization, energy accounting, and integration. Gate: correct constrained behavior without hidden numerical artifacts.
- P1Rigid-object interaction in simulation
Pushing, sliding, collisions, variable mass and friction. Gate: useful prediction on held-out combinations and contact graphs.
- P2One real tabletop system
Arm, RGB-D, wrist force/torque. Gate: measured contact prediction and lower adaptation cost after identification.
- P3Tactile contact
Insertion and constrained manipulation. Gate: compatible tactile representation improves closed-loop outcomes over matched vision/force baselines.
- P4Second embodiment
Another arm or gripper with the core frozen. Gate: calibration beats retraining in total cost while retaining earlier performance.
- P5VLA integration
Generalist policy for semantics and proposals, firmware for consequences. Gate: incremental benefit over the same VLA with equal sensing and compute.
- P6Industrial workflow
Repeated manipulation with changing parts. Gate: lower deployment cost at the required reliability across operating shifts.
P0 should be deliberately small. A successful pendulum validates code and assumptions; it cannot justify a claim about contact-rich robotics. The scientific center of gravity is P2–P4: real contact, identifiable grounding, and a core that survives a changed body.
23 / THE FIRST REAL EXPERIMENTOne arm, variable objects, and an honest comparison.
Proposed minimum credible experiment: use a Franka Panda or equivalent research arm with joint-state access, an RGB-D camera, and a calibrated wrist force/torque sensor. Begin with planar pushing and object repositioning on an instrumented, bounded tabletop. Tactile sensing and insertion are follow-ups, not prerequisites for the first result.
Make hidden physical changes measurable
Use objects with interchangeable internal weights and adjustable center of mass, several geometries and sizes, and replaceable surface materials. Measure reference mass, geometry, and friction under a declared procedure for evaluation; do not feed held-out parameter labels to the deployed model. Keep some appearances constant while changing weight so a visual shortcut cannot explain transfer.
Create disjoint splits for familiar objects, held-out combinations of known factors, and new parameter ranges. Reserve entire sessions and object configurations for evaluation. Use synchronized pose tracking as an evaluation reference—fiducials or another calibrated tracker are acceptable for the prototype—and report its error. Keep an oracle-state condition to diagnose perception, but make the main result use the same available sensors for both models.
Build two predictors that differ in the thing being tested
The unstructured candidate uses the same state estimator, object representation, probabilistic output family, training trajectories, and planner. It predicts transitions without the selected mechanical constraints. The structured candidate uses the same inputs with explicit geometry, an appropriate rigid-body update, a contact treatment, and a limited learned residual. Match capacity and compute as closely as possible. Then add ablations for graph structure, equivariance, and uncertainty separately.
For slow planar pushing, a quasi-static model—one that neglects inertia when its effects are small—may be the strongest structured candidate. Include it. Introduce faster motion and changes of velocity only if testing inertial structure. Otherwise a Hamiltonian network could appear ineffective simply because the experiment never required its contribution.
Implementation note: a minimal candidate worth building
Use explicit planar pose and velocity for each object; a robot tool node; shape features; a belief over friction and inertial properties; and actuator history. A recurrent encoder integrates visual pose estimates, joint states, and forces. A probabilistic transition ensemble predicts short action-conditioned trajectories, including contact mode and sensor-level force targets.
For the structured variant, separate free motion, contact response, and bounded residual prediction. Use an analytic or constrained contact update rather than assuming smooth energy dynamics will discover impacts. Estimate common physical parameters from a short calibration history. Preserve the same observation decoder and uncertainty family in the unstructured control.
Use a sampled MPC optimizer such as the cross-entropy method: repeatedly sample candidate command sequences, keep lower-cost candidates, and refit the sampling distribution. Give both predictors the same candidate budget and deadline. A separately tuned planner condition can test best practical performance, but must not replace the matched-planner comparison.
Run an adaptation ladder, then change the body
Evaluate before task-specific adaptation and after predeclared cumulative interaction budgets. Freeze the pretrained models between budget checkpoints. Count probing, failures, demonstrations, and autonomous exploration; report training time separately. Measure multi-step motion and force error, contact timing, parameter-shift transfer, calibration, MPC success, and the interaction needed to reach the same goal tolerance and reliability.
Next, change the arm or gripper while retaining the trained physical core. Allow only embodiment identification and limited adapter updates in the primary transfer condition. Compare against a freshly fitted predictor, a fine-tuned core, and the same policy baseline. A new gripper tests end-effector transfer; a genuinely different arm and control interface is a stronger embodiment test. Label each accurately.
A repeatable reduction in total new physical interaction at matched reliability, on held-out dynamics and then a changed embodiment, with the improvement traceable to specific structure. No fixed multiplicative gain is assumed.
24 / PRODUCT IMPLICATIONSOnly after the transfer result, ask where it pays.
Conditional product hypothesis: high-mix industrial manipulation could be a useful first application. The robot repeats a family of tasks while parts, fixtures, weights, or surface conditions change. A known workspace, measurable contact, robot telemetry, and explicit reliability targets make evaluation more tractable than a general household robot.
The value would be fewer engineering and operator hours to deploy a new variant, less downtime during changeover, or fewer expensive contact failures. These are testable operational outcomes. They must exceed the cost of extra sensors, calibration, inference hardware, integration, and maintaining the model.
A useful prediction layer might become a calibrated rollout service, an adaptation toolkit, or offline training infrastructure. The experiment should decide which form is justified. There is no market-size estimate or assumption that a technically elegant model becomes a standalone business.
25 / THE DATA FLYWHEELThe scarce asset would be grounded interaction history.
If the hypothesis works, repeated deployments could accumulate synchronized physical interaction data, rare failures, contact transitions, unusual materials, embodiment adapters, calibrated robot models, learned residuals, and cross-robot transfer records. A benchmark suite and a reliable account of deployment costs would be part of that asset.
The feedback loop is concrete: log prediction errors → identify a failure regime → collect bounded informative interactions → update a candidate → retest old and new conditions → deploy only after regression checks. More data is useful when it covers missing interactions and retains provenance; repeated nominal successes can reinforce the same blind spots.
Any defensibility would come from the combination of architecture, interaction data, embodiment grounding, evaluation, and deployment feedback. “We use Hamiltonian networks” is not a defensible claim to the underlying idea. Data rights, instrumentation quality, and reproducibility matter as much as collection volume.
26 / BOUNDARIESWhat Physical Firmware is not.
- A universal physics oracle, or a claim that Newtonian mechanics solves robotics.
- A conventional rigid-body simulator with a new name, though a candidate may include one.
- A replacement for a VLA, perception system, robot operating system, or low-level motor controller.
- Proof that architectural physics beats scaling or broad behavioral pretraining.
- A guarantee of cross-embodiment intelligence.
- A claim that all relevant manipulation can be expressed through Hamiltonian mechanics.
- An experimentally validated system or product.
World model ≠ policy. VLA ≠ simulator. System identification ≠ task learning. Uncertainty estimate ≠ calibrated uncertainty. Same task ≠ same dynamics. Keeping those distinctions visible is what makes the proposal testable.
27 / THE FINAL BETHow much knowledge can survive a change?
Robot learning is demonstrating that scale, heterogeneous data, and foundation models can produce increasingly general behavior. Physical Firmware asks a narrower question:
If a robot's predictive model is explicitly structured around geometry, constraints, contacts, uncertainty, and dynamical regularities, can more of what it learns survive the transition from one task, object, or robot to another?
The answer is not known. It is experimentally testable. That is the project.
Physical Firmware is a reusable, probabilistic, contact-aware and physically structured world model intended to reduce the amount of new physical experience required when a robot encounters a new task, environment or embodiment.
SOURCES / VERIFIED THROUGH 19 SEPTEMBER 2026References & further reading
Research lineage, not evidence that this proposal works. Peer-reviewed papers, preprints, and official reports are labeled separately. Dates refer to the cited publication or report, not a search engine's crawl date. No external benchmark numbers are reproduced.
- Physical Intelligence / Black et al. π0: A Vision-Language-Action Flow Model for General Robot Control.
- Physical Intelligence. π0.5: a Vision-Language-Action Model with Open-World Generalization.
- Physical Intelligence. π*0.6: a VLA That Learns From Experience.
- Physical Intelligence. π0.7: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities.
- Physical Intelligence. VLAs with Long and Short-Term Memory.
- Physical Intelligence. Precise Manipulation with Efficient Online RL.
- Open X-Embodiment Collaboration. Open X-Embodiment: Robotic Learning Datasets and RT-X Models.
- Chi et al. Diffusion Policy: Visuomotor Policy Learning via Action Diffusion.
- NVIDIA et al. GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.
- NVIDIA Research. GR00T N1.5: An Improved Open Foundation Model for Generalist Humanoid Robots.
- NVIDIA. Cosmos World Foundation Model Platform for Physical AI and Cosmos 3: Omnimodal World Models for Physical AI.
- Gemini Robotics Team / Google DeepMind. Gemini Robotics: Bringing AI into the Physical World; Hansen and Xu, Introducing Gemini Robotics ER 2.
- Google DeepMind. Genie 3: A new frontier for world models.
- Fangyuan Wang et al. World Models for Robotic Manipulation: A Survey.
- Sven Kirchner, Nils Purschke, and Alois Knoll. A survey of world models for physical AI with uncertainty representation and control.
- Zhiyuan Zhang et al. ContactWorld: What Representations Matter in Vision-Tactile World Models for Contact-Rich Manipulation.
- Caoliwen Wang et al. WorldContact: A Contact-Centric World Model for Scalable Robot Learning.
- Yunfan Lou et al. Dream-Tac: A Unified Tactile World Action Model for Contact-Rich Robot Manipulation.
- Danijar Hafner et al. Learning Latent Dynamics for Planning from Pixels.
- Kurtland Chua, Roberto Calandra, Rowan McAllister, and Sergey Levine. Deep Reinforcement Learning in a Handful of Trials using Probabilistic Dynamics Models.
- Michael Janner, Justin Fu, Marvin Zhang, and Sergey Levine. When to Trust Your Model: Model-Based Policy Optimization.
- Ricky T. Q. Chen, Yulia Rubanova, Jesse Bettencourt, and David Duvenaud. Neural Ordinary Differential Equations.
- Samuel Greydanus, Misko Dzamba, and Jason Yosinski. Hamiltonian Neural Networks.
- Miles Cranmer et al. Lagrangian Neural Networks.
- Thai Duong, Abdullah Altawaitan, Jason Stanley, and Nikolay Atanasov. Port-Hamiltonian Neural ODE Networks on Lie Groups for Robot Dynamics Learning and Control.
- Fabian J. Roth, Dominik K. Klein, Maximilian Kannapinn, Jan Peters, and Oliver Weeger. Stable Port-Hamiltonian Neural Networks.
- Alvaro Sanchez-Gonzalez et al. Learning to Simulate Complex Physics with Graph Networks.
- Víctor Garcia Satorras, Emiel Hoogeboom, and Max Welling. E(n) Equivariant Graph Neural Networks.
- David Stewart and Jeffrey C. Trinkle. An Implicit Time-Stepping Scheme for Rigid Body Dynamics with Coulomb Friction.
- Ernst Hairer, Christian Lubich, and Gerhard Wanner. Geometric Numerical Integration: Structure-Preserving Algorithms for Ordinary Differential Equations.
- Ashish Kumar, Zipeng Fu, Deepak Pathak, and Jitendra Malik. RMA: Rapid Motor Adaptation for Legged Robots.
- Yuanming Hu et al. DiffTaichi: Differentiable Programming for Physical Simulation.
- Rina Foygel Barber, Emmanuel J. Candès, Aaditya Ramdas, and Ryan J. Tibshirani. Conformal prediction beyond exchangeability.
- Shenghua Wan, Le Gan, and De-Chuan Zhan. Learning to Be Uncertain: Pre-training World Models with Horizon-Calibrated Uncertainty.
This article defines a research hypothesis, an architectural design space, and an evaluation program. Diagrams are explanatory schematics. The learning curves are hypothetical. No implemented Physical Firmware library, trained model, commercial deployment, or measured performance advantage is asserted.