MULTIMODAL IS THE NEW DEFAULT-YOUR WORKFLOWS AREN’T READY
YOUR WORKFLOWS AREN’T READY.
Your team is still typing into chat boxes. Frontier
models now take voice, screenshots, video, and
PDFs as a single input. The companies that retooled
their workflows are running 3x faster on knowledge
work.
Why it Matters
Two years ago, you talked to AI by typing. In
2026, you point your camera at a problem and
the AI sees it. You drop a thirty-page PDF and
a Zoom recording into the same conversation
and ask a single question. The interface didn’t
change overnight — most users still type —
but the back-end changed completely, and the
gap between what AI can do and what most
teams ask it to do is widening fast.
Problem
Mid-market teams adopted AI through chat
interfaces in 2023–24. Those workflows treat AI
as a text-in, text-out tool. The new generation
of frontier models — GPT-4o and successors,
Claude with vision, Gemini, the multimodal
Llama and Gemma variants — take any input
mix: text, images, screenshots, audio, video,
PDFs, spreadsheets. The teams still pasting
transcripts into a chat box are doing two hours
of work for every fifteen minutes the modern
workflow would have spent.
How
Three workflows show the productivity delta
most clearly. First, meeting follow-up: drop
the Zoom recording, the deck shown, and
the original brief into one conversation; the
AI delivers minutes, action items, and a draft
client follow-up in minutes. Second, document
review: a multi-document upload with mixed
PDFs and images yields a comparative analysis
no single-input workflow can match. Third,
visual troubleshooting: a field tech photographs
a problem, the AI cross-references with
documentation, and a fix is on screen before
the truck rolls. Companies that retooled report
3x faster knowledge work cycles on tasks that
benefit from multimodal input.
Why
This is one of the few AI capability shifts where
mid-market can move faster than enterprise.
The tools are off-the-shelf. The workflow
redesign doesn’t require IT investment — it
requires permission for teams to experiment
with how they use what they already pay
for. The companies that won the first AI
productivity wave were the ones that retooled
their workflows. The same opportunity is open
right now for the multimodal wave, and the
window is shorter.
Action
Three workflow redesigns to run this quarter.
First, replace any process where someone
transcribes a meeting before doing analysis —
the transcription step is now part of the AI call.
Second, replace any process where someone
manually compares documents — multi-
document AI workflows do this in one pass.
Third, audit any field service or operations
process where a human describes a visual
problem to a remote expert — multimodal AI
with the right setup can short-circuit that loop
entirely.
Pushback
“Our team won’t change how they work.” They
will if the productivity case is concrete and the
new workflow is faster than the old one on day
one. The trick is picking the right workflows to
start with — ones with visible time savings, not
abstract benefits.
KEY TAKEAWAYS
• Frontier models now accept any input
mix: text, images, audio, video, PDFs,
spreadsheets in a single conversation.
• Companies retooling workflows
report 3x faster knowledge work
cycles on multimodal-friendly tasks.
• Three high-leverage targets:
meeting-to-deliverable workflows,
multi-document review, and visual
troubleshooting.
• This is a workflow redesign play, not
an IT investment — mid-market can
move faster than enterprise here.















