Skip to content
LearnSign up

Learn / The Academy · 2 min read

Vision-language models

AI that can look at a picture and talk about it

The short answer

A vision-language model is AI that can look at a picture or video and talk about it in plain words. Show it a photo of a restroom and it can tell you the trash is full and the mirror is streaked. It sees and it talks. On its own, it can't move anything.

First-person view of a restroom mid-clean: a full trash can, a streaked mirror and a wet floor sign.

Eyes plus words

"Vision" means it takes in pictures. "Language" means it reads and writes words. Put them together and you get a program you can ask, "What still needs doing in this room?"

You may already use one. Some phone apps let you snap a photo of a label and ask what the chemical is for. That is a vision-language model at work.

What it gets right, and what it misses

These models are good at naming things. Mop bucket, wet floor sign, paper towel dispenser: most will get those.

They are weaker on the work itself. Is that stall done, or just wiped once? Did the cleaner hit the handle and the flush lever? Your supervisors can tell in two seconds on a walk-through. A model needs to see a lot of real jobs, start to finish, to learn that.

Where it fits in a robot

Some robots use a vision-language model as the part that thinks. Figure says the top layer of its Helix system is one of these models. It looks at the scene and works out the goal about 7 to 9 times a second. A faster part then moves the hands 200 times a second.

So the model plans and a second part carries it out. The next lesson covers models that do both.

Why this matters to your crew

A model learns what "done" looks like from video where the job actually gets done. When your crew records a full restroom, from the first trash pull to the last mirror wipe, that clip teaches the whole order of the work. Short, staged clips can't teach that.

What other owners ask

Is a vision-language model the same as a chatbot?

It is a close cousin. A chatbot works with words. A vision-language model also takes in pictures or video.

Can it tell who is in a video?

That depends on the model. It is one reason we ask your crew to keep faces, papers and screens out of the shot and tell us fast if something private gets recorded.

Where to read more

These are public examples from other companies. They are not Colby partners or customers.

Next lessonVision-language-action models

Still have questions?

Talk to our team.

Call us and ask anything about pay, your clients or your crew. Or sign up and we will call you.