Physical AI means models that see, talk, and act on robots — usually a vision-language-action (VLA) stack plus a simulator.
What’s actually hard
- Sim-to-real gap (friction, lighting, contact)
- Safety and kill switches
- Fleet economics, not a single demo video
How to evaluate vendors
- Ask for a reproducible task (pick-and-place under lighting change), not a showreel
- Prefer stacks with open eval harnesses and logged failures
- Separate foundation model claims from robot OS / fleet claims
Hosted LLM / video boards on Models & Makers track adjacent multimodal pieces (e.g. video generators). Robotics SKUs are more fragmented — treat marketing parameter counts carefully and re-verify before procurement.