Skip to main content

Multimodal Prompt Engineering: A Practical Guide to Images, PDFs, Audio & Video

Calculating read time…

Multimodal prompt engineering is the practice of deciding what an image, PDF, audio clip, or video actually becomes once it's converted into the tokens a model reasons over — and no two providers convert them the same way. 🎛️

Treat every input format as interchangeable and a system breaks in quiet, expensive ways: a chart gets summarized from its caption instead of its bars because the image was too small to read, a 40-minute recording gets billed as if every silent second still cost tokens, or a workflow gets built around a modality — raw video, raw audio — that the model on the other end of the API call was never actually trained to accept. Knowing what happens to each format on the way in is the difference between a demo that worked once and a pipeline that holds up in production. 📉

Diagram showing images, PDFs, audio, and video each passing through a different converter before becoming a shared token stream inside one context window

🔀 Quick Comparison

Modality How It Reaches the Model Cost Driver Where Support Varies Most
Image Split into pixel patches and converted into a countable set of tokens Resolution — bigger image, more tokens billed Widely supported across major providers, but exact token math per pixel differs
PDF Each page rendered as an image and its text extracted, together Page count and file size, plus whatever image cost each page adds Page and size limits vary by vendor and model
Audio Understood natively by some models; transcribed to text first for others Duration — longer clips cost more regardless of how much is silence The biggest capability gap in this whole post — check per model, not per vendor
Video Sampled into frames at a fixed rate plus a separate audio track, both timestamped Duration and resolution together — long, high-res video is the most expensive input format here Native video understanding is not universal — some models need it pre-processed into frames

1. What "Multimodal" Actually Changes About a Prompt

🧒 Kid analogy: imagine describing your day to a friend three different ways — over a text message, over a phone call, or by handing them your phone with a video already playing. Each way carries different information (the text message can't carry your tone of voice; the phone call can't show them your friend's face). A prompt with an image, a PDF, audio, or video attached is the same idea: you're not just changing what you say, you're changing which "channel" the model receives it through, and each channel has its own rules for what gets through clearly and what gets lost.

Mechanically, a multimodal prompt is still just a sequence of content blocks sent to the model — a text block, an image block, a document block — assembled into the same context window covered by ordinary prompt-engineering fundamentals like structure, ordering, and token budget. What's genuinely new is that each non-text block has to be converted into something token-like before the model can reason over it at all, and that conversion step is where providers diverge the most: how an image becomes tokens, whether a PDF's chart and its caption are both preserved, whether audio is understood directly or has to be transcribed first, and whether video is a first-class input or something you have to break into frames yourself.

🎯 Use this when: starting any new multimodal integration — before writing a prompt, confirm which modalities the specific model you're calling actually accepts natively, because that answer is different for every model, not just every vendor.

2. Working With Images

🧒 Kid analogy: if you photograph your homework from across the room, it looks like a tiny white rectangle — technically a picture of the homework, but useless for actually reading it. Get close enough and the words become clear, but the photo file also gets bigger. Sending an image to a model is the same trade-off: too small or blurry and the model can't read what matters; unnecessarily huge and you're paying for detail nobody needed.

Real-world example: Anthropic's vision documentation gives a specific, testable structural recommendation: while images placed after text or interleaved with it still work, they explicitly recommend an image-then-text ordering for best results, and note that a single request can include multiple images — up to 20 through claude.ai and up to 100 through direct API calls — with three delivery mechanisms available: base64-encoded data inline, a URL reference, or a one-time upload through the Files API that can then be reused across many requests without re-sending the bytes each time. Separately, OpenAI's own developer documentation describes images being converted into billable input tokens through a resolution- and patch-based system, where an image's pixel dimensions map directly to a countable number of tokens that also count against per-minute rate limits — making resolution an explicit, calculable line item rather than an incidental detail.

Diagram showing a small low-resolution image producing few token patches and a full-resolution version of the same image producing many more

Two structural facts fall directly out of that: first, image position in the prompt is a real design choice with documented guidance behind it, not a cosmetic one — check the specific model's current docs rather than assuming a universal rule, since Anthropic's stated preference doesn't automatically apply to every other vendor's models. Second, resolution is a lever with two opposite effects at once: turning it up can raise accuracy on tasks like reading small text or distinguishing similar-looking objects, while also directly raising token cost — so the right move is usually the smallest resolution that still lets the model answer correctly, not the largest one available.

✅ Worked example: A receipt-scanning tool for expense reports sends one moderately-sized image per receipt, in an image-then-text order — the photo first, then "Extract the merchant name, date, and total as JSON" — because the numbers on a receipt are usually large enough to read at a modest resolution, keeping cost predictable across thousands of receipts a month.

💡 Contrasting case: The same tool pointed at dense engineering schematics with tiny embedded part numbers needs meaningfully higher resolution to read those labels correctly — applying the receipt tool's low-resolution default here wouldn't fail loudly, it would just silently misread part numbers, which is a harder bug to catch than an outright error.

🎯 Use this when: the image contains small text, fine detail, or many similar-looking regions — that's the signal to test higher resolutions rather than defaulting to the cheapest setting.

3. Working With PDFs

🧒 Kid analogy: a picture book works two ways at once — you can look at the pictures to understand the story, or read the words to know exactly what happened. A good reader uses both together, because the picture might show a mood the words don't spell out, and the words might specify a detail the picture only hints at. A well-handled PDF gets read the same double way: as a picture (for charts, layout, and visual structure) and as words (for exact wording you can quote).

Real-world example: Anthropic's PDF support documentation describes exactly this dual-extraction approach: each page of a submitted PDF is processed both as a rendered image and for its underlying text, so a single request can answer questions that depend on a chart's visual shape and questions that depend on a specific sentence's exact wording, using the same document. The same documentation sets concrete operating limits — a maximum file size and a maximum page count per document — and describes the same three delivery paths available for images (a hosted URL, base64-encoded data, or a Files API upload), which also means a PDF benefits from prompt caching the same way a repeated image or text block would when the same document gets queried multiple times in a row.

Diagram showing one PDF page being read twice: once rendered as an image for its chart and layout, and once extracted as text for exact quoting

The practical takeaway is that a PDF prompt should ask questions that actually make use of both extraction paths, not just one. "Summarize this report" mostly exercises the text path. "What does the trend in the Q3 chart on page 4 suggest, and what specific sentence in the surrounding text explains it?" exercises both — and is exactly the kind of query where a text-only pipeline (one that ran the PDF through a plain text extractor before it ever reached the model) would have already thrown the chart away.

✅ Worked example: For a 60-page financial report, instead of asking one broad question over the whole document, a page-range-scoped question — "Summarize pages 14 to 22, and list every numeric figure that appears in a table on those pages" — grounds the answer in an identifiable, checkable section rather than letting the model range freely across 60 pages of mixed text and charts.

💡 Key warning: A scanned PDF that's actually just a stack of photographed pages has no separate machine-readable text layer to extract — the "read it as text" path returns nothing useful, and the system is quietly relying entirely on the image path, with all of the resolution trade-offs from Section 2 now applying to every single page.

🎯 Use this when: the document mixes narrative text with charts, tables, or scanned pages — that combination is exactly what a text-only extraction pipeline handles worst.

4. Working With Audio

🧒 Kid analogy: if a friend hums you a song over the phone, you can often guess the tune from the melody alone — the pitch, the rhythm, the pauses. If they instead text you the lyrics, you get the exact words but lose all of that. Some models can genuinely "hear" audio the way your friend on the phone hears the hum — pitch, tone, background noise and all. Others only ever see a written transcript, like the text message, which means anything that isn't captured in words simply isn't there for them.

Real-world example: Google's Gemini API documentation describes audio as a natively understood input: the model can analyze an uploaded audio file directly and answer questions not just about what was said, but about non-speech sound in the clip — their documentation specifically calls out things like birdsong or sirens as examples of sound the model can identify without a word of it ever being spoken. This is a meaningfully different capability from a text-only reasoning model paired with a separate transcription step, which is the pattern used where native audio understanding isn't available: an automatic speech recognition system (a Whisper-style transcriber, for instance) converts the audio to text first, and only that transcript reaches the language model — at which point tone of voice, background sound, and anything conveyed by silence or emphasis is already gone, because it was never in the transcript to begin with.

This matters concretely for anything checking how something was said rather than only what was said: a call-quality review that needs to flag a frustrated tone, a voicemail triage system that should notice a caller talking over someone else, or a podcast clip search that needs to find "the part with applause" rather than a specific sentence. Text-only pipelines built on a transcript alone cannot see any of that, by construction — no amount of clever prompting recovers information the transcription step already discarded.

🎯 Use this when: the task depends on tone, non-speech sound, or timing between speakers — that's the signal to check for a model with genuine native audio understanding rather than defaulting to transcribe-then-reason.

5. Working With Video

🧒 Kid analogy: a flip-book isn't a smooth movie — it's a stack of still drawings that look like motion when you flip through them fast enough. If a page is missing, you don't get a stutter, you just never see what happened on that page at all. That's a closer picture of how a model actually experiences a video than "watching" it: it's handed a stack of still frames, sampled at some fixed rate, plus a separate track for the sound — and anything that happened entirely between two sampled frames is simply not part of the flip-book.

Real-world example: Google's Gemini documentation is specific about this mechanism rather than leaving it abstract: when a video is uploaded through their Files API, it's stored and processed at a fixed sampling rate — one frame per second for the visual track, alongside the audio track sampled continuously — with a timestamp attached roughly every second, which is what lets the model answer a question like "what happens at the 0:45 mark" instead of only describing the video in general terms. Models with very large context windows can process significantly longer videos this way, with the documented ceiling depending on resolution — roughly an hour at standard resolution or several hours at a lower one, according to their published guidance, figures which are worth re-checking against current docs before relying on them, since sampling rates and duration limits are exactly the kind of detail vendors adjust between model versions.

Timeline diagram showing a video's visual track sampled once per second as still frames, alongside a continuously sampled audio track, both carrying timestamps

Not every model accepts video as a first-class input at all. Where native video understanding isn't available, the older and still-common workaround is to do the sampling yourself before the model ever sees it: extract still frames from the video at a chosen interval, extract the audio separately and transcribe it, and send the frames as a sequence of images alongside the transcript as text — manually reconstructing the same frame-plus-audio-track pattern a model with native video support would otherwise handle internally. The trade-off is direct: you have full control over the sampling rate (useful if 1 frame per second would miss something fast-moving that matters), but you're also doing, and paying for, extra preprocessing work that a natively video-capable model does for you.

✅ Worked example: here's the shape of the manual frame-extraction workaround, walked through before the code. It takes a local video file and a chosen sampling interval, then does three things: (1) uses a video-processing library to step through the video and pull out one still image every interval_seconds, recording the timestamp each frame came from; (2) runs the video's separate audio track through a speech-to-text transcriber to get a plain-text transcript; and (3) builds a single prompt that lists each frame as an image labeled with its timestamp, followed by the full transcript as text, so the model can cross-reference "what was said" against "what was on screen" at roughly the same moment.

frames = extract_frames(video_path, interval_seconds=2)
# frames is a list of (timestamp, image_bytes) pairs
transcript = transcribe_audio(extract_audio(video_path))

content = []
for timestamp, image_bytes in frames:
    content.append({"type": "text", "text": f"Frame at {timestamp}s:"})
    content.append({"type": "image", "source": image_bytes})
content.append({"type": "text", "text": f"Full transcript:\n{transcript}"})

💡 Key warning: A 2-second sampling interval will simply never catch a quarter-second visual event — a warning light that flashes once, a fast hand gesture. If the task depends on something that brief, either shorten the interval around the section that matters or use a model with native video support that samples more finely by default.

🎯 Use this when: choosing between native video support and the manual frame-extraction workaround — the deciding question is almost always "does anything important in this video happen faster than my sampling rate can catch?"

6. Hands-On Lab: Feel the Modality Gap Yourself

This lab uses one disposable photo and one disposable short recording — nothing you'll need to clean up, and no code required for the first half.

1
Take a photo of a page of dense text — a book, a printed article, anything with small print — but take it from far enough away that the text looks blurry and small in the frame. Upload it to a chat interface and ask "What does the third paragraph say, word for word?"
Expect to see: the model either declines to guess precisely, guesses wrong, or asks you for a clearer image — this is resolution insufficiency in action, the same trade-off from Section 2's diagram.
2
Take the photo again, this time close enough that the text is sharp and legible to your own eye, and ask the same exact question.
Expect to see: an accurate quote this time. Troubleshooting: if it's still wrong, check for glare, an angled shot, or genuinely tiny font — none of those are fixed by prompt wording, only by a better photo.
3
Record 15 seconds of yourself reading a sentence aloud twice — once completely flat and monotone, once clearly annoyed or excited. In a chat interface that accepts audio directly (check which model you're using — this is exactly the native-vs-transcript gap from Section 4), upload the clip and ask "Does the speaker sound frustrated in this recording?"
Expect to see: a model with genuine audio understanding should notice the tonal difference between your two readings, even though the words are identical.
4
Now type out a plain-text transcript of just the words you said — no tone markers, nothing about how it sounded — and ask a text-only model the same question: "Does the speaker sound frustrated?" using only that transcript.
Expect to see: a much less confident or outright wrong answer, since "I am so excited about this" and "I am so excited about this" (said flatly) are identical strings on the page — this is the information loss from Section 4's warning, now something you've triggered on purpose rather than discovered by accident in production.

Bridging to production: step 4 was you manually doing what a transcribe-then-reason pipeline does automatically at scale — and losing the same information a call-center sentiment pipeline would lose if it relied on transcripts alone. Steps 1–2 are exactly the resolution trade-off a document-scanning pipeline has to tune once, for thousands of images, instead of once, for one photo.

7. Rolling This Out at Enterprise Scale

Multimodal inputs raise the stakes on several of the same governance questions that apply to any prompt-engineering rollout, plus a few that are specific to non-text data:

  • Data governance for sensitive media: images, audio, and video routinely contain personally identifying information — faces, voices, license plates, handwriting — that a plain text prompt usually doesn't. Access to whatever pipeline uploads this media needs the same scrutiny as access to the underlying source system (a document store, a call-recording archive), not the lighter review a text prompt template might get.
  • Cost governance by modality: track spend separately for images, PDFs, audio, and video rather than lumping them into one token-cost number — a single long, high-resolution video can cost dramatically more than a thousand short text queries, and that imbalance is invisible in an aggregate dashboard.
  • Versioning across model upgrades: modality support is one of the least stable parts of a model's capability surface — a new model version can change resolution handling, add or drop native audio or video support, or change file-size and duration limits entirely. Pin and re-validate multimodal prompts against a regression suite before rolling out a model upgrade, the same discipline recommended for any prompt change.
  • Observability for silent degradation: a resolution setting, sampling rate, or transcription step that was fine for last quarter's inputs can quietly stop being fine as inputs drift — thicker documents, longer recordings, lower-quality scans. Track accuracy on a held-out sample per modality, not just overall request success rate, since a request can "succeed" (return a response) while quietly answering from too little information.
  • Fallback design: build an explicit fallback for when a modality isn't supported by the model in use — routing video to a frame-extraction pipeline, or flagging an audio file for transcription — rather than discovering the gap when a request errors out or, worse, silently ignores the attached media.

🎯 Use this when: multimodal inputs are moving from a prototype into a system handling real customer or employee data — that's the point where the governance questions above stop being optional.

8. Common Mistakes

  • Assuming every model in a product lineup supports the same modalities. Vision, PDF, audio, and video support genuinely differ model to model, even within one provider's lineup — a prompt built and tested against one model can silently fail, or silently ignore an attachment, on another.
  • Sending maximum resolution "just in case." Without testing whether a lower resolution already answers correctly, teams default to the largest image size available, paying for detail the task never used — the inverse mistake of Section 2's schematic example, applied indiscriminately instead of selectively.
  • Treating a transcript as equivalent to the audio. A transcript captures words, not tone, pacing, overlapping speech, or non-speech sound — building a pipeline as if a transcript were a lossless stand-in for the audio produces confident-sounding answers to questions the transcript was never able to answer.
  • Ignoring the sampling rate on video. Assuming a video "was watched" rather than sampled at a fixed rate leads to confusion when the model misses a brief event — the fix is adjusting the sampling interval or switching to native video support, not rewording the prompt.
  • Skipping the text-extraction path on documents with real prose. Relying only on a PDF's rendered-image path to answer a question that needs an exact quote risks the model paraphrasing instead of quoting precisely, when the text layer would have given an exact match.
  • Not testing on the messiest real-world inputs. A demo built on a clean, well-lit photo, a studio-quality recording, or a perfectly formatted PDF says little about how the same pipeline handles a phone photo at an angle, a noisy voicemail, or a scanned document with handwritten annotations — exactly the inputs production traffic actually contains.

❓ FAQ

Does every model that accepts images also accept PDFs, audio, and video?

No. Image support is the most widely available multimodal capability across major providers, but PDF, audio, and video support vary considerably from model to model, even within the same provider's lineup. Always check the specific model's current documentation rather than assuming capability carries over from one modality to another.

Should I always use the highest image resolution available for the most accurate answer?

Not necessarily. Higher resolution can improve accuracy on tasks needing fine detail, but it also increases token cost directly. The better approach is testing the smallest resolution that still answers correctly for your specific task, rather than defaulting to the maximum.

If a model transcribes audio instead of understanding it natively, is that a real limitation?

Yes, for tasks that depend on tone, non-speech sound, or timing. A transcript preserves words but discards how they were said, so questions about sentiment, emphasis, or background sound can't be reliably answered from a transcript alone, no matter how the prompt is worded.

Why would a model miss something that clearly happens in a video I sent it?

Video is processed as a series of sampled still frames plus a separate audio track, not watched continuously. An event that happens entirely between two sampled frames — faster than the sampling rate — simply isn't captured in any frame the model receives.

Is a scanned PDF handled the same way as a native, text-based PDF?

Not entirely. A native PDF has both a visual layout and an underlying text layer that can be extracted precisely. A scanned PDF (effectively a stack of photographed pages) has no separate text layer, so the system depends entirely on the image-based path, making resolution and image quality far more important for scanned documents.

🔗 References & Further Reading

Official / primary sources relied on for this post:

All product and company names (Anthropic, Claude, Google, Gemini, OpenAI, GPT, and others referenced) are trademarks of their respective owners. 

📝 Summary

  • Every non-text modality gets converted into tokens differently, and that conversion step is where provider and model capabilities diverge the most.
  • Images become billable tokens based on resolution — bigger isn't automatically better, since it's a cost decision as much as a quality one.
  • A well-handled PDF is read twice: once as an image (for layout and charts) and once as text (for exact wording) — relying on only one path loses real information.
  • Audio is either understood natively — including tone and non-speech sound — or reduced to a transcript first, and that distinction determines what questions can actually be answered.
  • Video is sampled into timestamped frames plus an audio track, not watched continuously — fast events between samples are genuinely invisible to the model.
  • The hands-on lab lets you personally trigger a resolution failure and an audio-transcript information loss, rather than just reading about them.
  • At enterprise scale, multimodal inputs raise the stakes on data governance, per-modality cost tracking, and re-validation across model upgrades.
  • The most common mistakes all come from assuming one modality's rules (or one model's capabilities) automatically apply to another.

That's multimodal prompt engineering end to end — from the flip-book behind every video to the two-track reading behind every PDF. Go check what your specific model actually accepts before you build around an assumption. 🎬

Comments