Image search no longer depends only on captions: AI Search now embeds image pixels directly with Qwen3-VL-embeddingA list of numbers representing a piece of text's meaning, so that similar meanings end up numerically close and can be searched.Full definition →, so queries can match visual detail a caption left out, such as texture, markings or layout. Captions are still kept for text understanding.
Text-only embedding models can still take image queries: AI Search converts the query image to text with ToMarkdown and searches on that caption. Models with native image support embed the query image into the same vector space as indexed images and text.
Retrieval runs vector and keyword search in parallel, then fuses and optionally reranks the results. Queries can be rewritten first, and the top chunks are either returned directly or passed to a generation model that writes an answer.
The file size limit for text files and PDFs rose from 4 MiB to 10 MiB. Scanned PDFs with no extractable text can use optional OCR, which reads each page before chunkingSplitting documents into passages small enough to embed and retrieve individually — the step that quietly determines whether retrieval works at all.Full definition → and is billed as image processing ingestion tokenThe chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.Full definition →.
From November 1, 2026, Cloudflare lists $0.75 per 1M ingestion tokens (+$0.50 for images), $2.00 per GB-month stored, and $0.75 or $0.10 per 1k semantic or full-text queries. Parsing, chunking, Workers AI embedding and rerankingA second pass that re-scores retrieved candidates by reading each one against the query, fixing the ordering that fast vector search got approximately right.Full definition → are included.
Terms in this piece · Glossary
embedding — A list of numbers representing a piece of text's meaning, so that similar meanings end up numerically close and can be searched.
chunking — Splitting documents into passages small enough to embed and retrieve individually — the step that quietly determines whether retrieval works at all.
token — The chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.
reranking — A second pass that re-scores retrieved candidates by reading each one against the query, fixing the ordering that fast vector search got approximately right.