Skip to content
Live
Loading the latest AI news…

Glossary · How models work

Multimodal AI

Multimodal AI can take in or produce more than one kind of information — text, images, audio, video — in the same conversation. You can show it a photo and ask a question about it out loud.

In one line, for a 12-year-old

Multimodal AI can read, look, listen and talk, not just type.

An example

Photographing a confusing parking sign and asking your assistant, "Can I park here at 6 p.m. on a Tuesday?"

Why it matters to people

Photos and voice carry more personal information than text: faces, locations, background details. Check what's in the frame before you share.