Multimodal AI
Multimodal AI can take in or produce more than one kind of information — text, images, audio, video — in the same conversation. You can show it a photo and ask a question about it out loud.
In one line, for a 12-year-old
Multimodal AI can read, look, listen and talk, not just type.
An example
Photographing a confusing parking sign and asking your assistant, "Can I park here at 6 p.m. on a Tuesday?"
Why it matters to people
Photos and voice carry more personal information than text: faces, locations, background details. Check what's in the frame before you share.