Skip to content
LearnSign up

Learn / The Academy · 2 min read

Vision-language-action models

AI that sees, understands and then moves

The short answer

A vision-language-action model, or VLA, is AI that looks at the room, reads an instruction and then moves a robot's arms and hands. You say "empty that trash can" and it works out the motions. It is the kind of model many robot teams are building now.

First-person view of a cleaner pulling a full trash liner out of a break-room can with both hands.

Adding the "action"

The last lesson covered models that see and talk. A VLA adds a third job: it sends moves to the robot's motors. Reach, grip, lift, turn.

Google DeepMind calls its Gemini Robotics model a VLA that "converts vision and language input into motor control." Figure calls its Helix system a VLA too. We name them only as examples of the field.

One trash can, many small moves

Tell a VLA "empty the trash can under the sink." It has to find the can. Then reach in, pinch the liner edge, and lift without tearing it. Then tie it, drop in a new liner, and fold the edge over the rim.

Your crew does that in 20 seconds without thinking. For a robot, each of those steps is a separate hand problem. The model learns them by seeing people do it many times.

Where they still struggle

Many robot demos happen on neat tables in labs. Real buildings are messy. A liner stuck to the can with old coffee. A mop bucket in a tight janitor closet. A soap dispenser that jams.

Those odd moments are exactly what real cleaning video shows. They are hard to stage and easy to catch on a normal Tuesday night shift.

Why this matters to your crew

A VLA copies hand motions, so it needs to see hands clearly doing the real task. Your crew's hours help most when the phone is angled down, both hands are in frame and the work is the actual job at a normal pace. The messy parts of a real shift are useful too, so there's no need to tidy up the building before you record.

What other owners ask

Is a VLA the same as a robot?

No. The VLA is the program. The robot is the body it runs on.

Does a VLA learn from our video directly?

Teams train these models on many kinds of data, often including first-person video of people. We do not say which team uses any given hour.

Where to read more

These are public examples from other companies. They are not Colby partners or customers.

Next lessonWorld models

Still have questions?

Talk to our team.

Call us and ask anything about pay, your clients or your crew. Or sign up and we will call you.