Inverity
SEO & Discovery

Image SEO in the AI-Search Era: The New Citation Currency

Date Published

An image tile with alt text and schema metadata chips being cited by two abstract AI answer panels, with citation arrows pointing from the image to a thumbnail and source link inside each answer

Most image SEO advice is a fossil. Compress your files, write your alt text, submit a sitemap: the checklist has barely moved since 2015, while the thing it optimizes for, the ranked blue-link results page, is being dismantled in public. By early 2026, 68.01% of US Google searches ended without any click at all, up from 60.45% in 2024 (SparkToro, Datos clickstream analysis, 2026).

Here is the argument of this post: as text links stop getting clicked, images are becoming the citation currency of AI search. They are the surviving surface where a brand is visibly credited, visually differentiated, and still clickable inside an AI-generated answer. And the industry is spectacularly unprepared for this, because nearly half of all web images are invisible to the text-first retrieval systems that decide what gets cited.

That is a thesis, not a settled fact, so we will show the data on both sides, including Google's rebuttal.

TL;DR >- When an AI summary appears, users click a traditional result on only 8% of visits, versus 15% without one; links inside the summary get clicked on just 1% of visits (Pew Research Center, 2025)- 68.01% of US Google searches now end without a click, and only 276 of every 1,000 searches produce an open-web click, down from 374 in 2024 (SparkToro/Datos, 2026)- Google's AI Mode now runs "visual search fan-out," analyzing primary and secondary objects in images and surfacing shoppable results from a 50B+ listing Shopping Graph (Google, Sept 2025)- Only 55% of web images carry non-blank alt text, which means nearly half the web's images are unreadable to text-first AI retrieval (Web Almanac 2024, HTTP Archive)- Google disputes the click-suppression findings and says AI Overviews reach 1.5B monthly users (Search Engine Land, 2025)

Why are images becoming the citation currency of AI search?

Because three trends converge on one conclusion: text links are losing clicks, AI answers are absorbing them, and images are the element AI surfaces still render as attributable, clickable objects. Pew measured traditional-result clicks falling from 15% to 8% of visits when an AI summary is present (Pew Research Center, 2025). Whatever survives that squeeze is where visibility now lives.

Consider what an AI answer actually looks like in 2026. The text is synthesized: your carefully written paragraph gets paraphrased into an unattributed sentence, with a citation link that Pew found gets clicked on 1% of visits. The images are not synthesized in the same way. When an AI surface shows a product photo, a chart, or a diagram, it shows your asset, sourced from your domain, often as the visual anchor of the whole answer.

Google made this concrete on September 30, 2025, when it launched visual results in AI Mode using what it calls "visual search fan-out": Gemini analyzes the primary subject of an image and its secondary objects, then assembles shoppable, visual answers drawing on a Shopping Graph of more than 50 billion product listings (Google, September 2025). The answer engine is not just quoting text anymore. It is actively decomposing and re-surfacing images.

The strategic read: in a results page, your image decorated your link. In an AI answer, your image often is your presence. That inversion is the whole game.

The click suppression, quantified

Bad, and measurably worse where AI summaries appear, though the size of the effect is contested. Pew's analysis of 68,879 real-user queries from March 2025 found users clicked a traditional result on 8% of visits with an AI summary present versus 15% without (Pew Research Center, 2025). Nearly half the clicks, gone.

The same study found 58% of panelists encountered at least one AI summary during the study month, and sessions ended entirely after an AI-summary page 26% of the time, versus 16% for traditional results pages. Users are not just clicking less; they are leaving satisfied, or at least leaving.

Zoom out from individual queries to the whole click economy and the picture holds. SparkToro's clickstream work puts US zero-click searches at 68.01% in early 2026, up from 60.45% in 2024. Of the clicks that remain, roughly 27% stay inside Google's own properties: AI Mode, YouTube, Maps, and, notably for this post, Google Images. Only 276 of every 1,000 US searches now produce a click to the open web, down from 374 in 2024 (SparkToro/Datos, 2026, with 2024 baseline from the prior study).

Line chart: the zero-click share of US Google searches rose from 60.45% in 2024 to 68.01% in early 2026

Zero-click searches climbed 7.56 percentage points in two years (SparkToro/Datos clickstream analysis, 2026).

The counterpoint: Google disputes the data

Fairness requires the other side. Google publicly disputes Pew's findings, calling the methodology flawed, and says AI Overviews now reach 1.5 billion monthly users while driving quality traffic to the web (Search Engine Land, 2025). Google argues query volume is growing and that aggregate clicks tell a different story than per-query rates.

Both things can be true: total clicks can hold up while the probability of any given page earning a click collapses. For an individual publisher, the per-query number is the one that pays the bills. We would also note that third-party trigger-rate studies for AI Overviews vary wildly, from 13% to 48% depending on the keyword set, so treat any single trigger-rate figure with suspicion. The direction of every independent dataset, however, points the same way.

How do AI systems actually "see" your images?

Mostly, they don't. They read text about your images. ChatGPT, Perplexity, and Google's retrieval layers are text-first systems: alt text, surrounding captions and headings, filenames, and structured data are the primary handles they have on an image at retrieval time. Which makes this figure alarming: only 55% of web images carry non-blank alt text (Web Almanac 2024 Media, HTTP Archive).

For twenty years, missing alt text was framed as an accessibility failure, which it is, and that framing let most businesses deprioritize it. The reframe that matters now: 45% of the web's images are functionally invisible to the systems deciding what gets cited in AI answers. An image without text handles cannot be retrieved for a query, cannot be matched to intent, and cannot become the visual anchor of an answer. The accessibility argument failed to move budgets for two decades. The AI-visibility argument might take two quarters.

The retrieval behaviors differ by engine, and the differences are strategically useful. Perplexity retrieves and cites on essentially every answer, so well-labeled images on topically strong pages have frequent chances to surface. ChatGPT often answers from its training data without retrieving at all, and when it does retrieve, it skews toward high-authority sources. Citation patterns diverge substantially between engines: Profound's analysis of 680 million citations documents distinct source preferences per platform rather than one shared canon (Profound, 2025). You are optimizing for several different librarians, not one.

There is a deeper question lurking here: when multimodal models do look at pixels, what do they actually evaluate? We've dug into that in how AI models judge visual quality, and the short version is that machine judgment and human judgment diverge in ways that matter for anyone optimizing images for machine consumption.

Which image SEO levers does Google officially confirm?

Google's documented levers are unglamorous and mostly ignored: descriptive alt text, meaningful filenames, image sitemaps, structured data (ImageObject, Product, Article), IPTC embedded metadata, and fast loading with responsive markup as quality signals (Google Search Central, image best practices). Nothing exotic. The edge comes from the fact that almost nobody does all of them.

Two of these deserve elaboration because they connect directly to AI surfaces.

Structured data is the machine-readable claim about your image. For product images feeding rich results and, increasingly, AI Mode's shoppable answers, Google publishes hard requirements:

Requirement

Google's specification

Apparel product images

Minimum 250 x 250 px

Other product images

Minimum 100 x 100 px

Maximum image size

64 megapixels

Recommended width for rich results

1200 px or wider

Licensable badge eligibility

Structured data or embedded IPTC metadata

Source: Google Search Central, Google Images best practices (current as of July 2026).

That last row is underappreciated: either structured data or embedded IPTC photo metadata qualifies an image for the Licensable badge. IPTC survives inside the file itself, which matters in an era when your images travel into answer engines without their surrounding page.

Performance is image SEO, full stop. Google lists fast loading and responsive markup as quality signals, and on most pages the LCP element is an image. An oversized hero image is simultaneously a Core Web Vitals problem, a ranking drag, and a worse candidate for image surfaces. The playbook for that overlap is in our Core Web Vitals image guide, and the compression judgment behind it, how far you can compress before quality signals suffer, is the subject of the complete guide to perceptual media optimization.

What to do differently now

Treat every image as a potential citation, which changes priorities in five concrete ways. The teams that win AI-surface visibility over the next two years will be the ones that made images retrievable, verifiable, and original while competitors were still A/B testing title tags.

  1. Close the alt-text gap as an AI-visibility project, not a compliance chore. Audit coverage across your top templates. If you're near the 55% web baseline (Web Almanac 2024), roughly half your visual inventory is unretrievable. Write alt text that describes the image's subject and context in a full sentence, the same text a retrieval system would need to match a query.
  2. Ship structured data on every product and article image. Meet the pixel minimums above, provide multiple aspect ratios at 1200px+, and add ImageObject markup with license fields. AI Mode's shoppable fan-out is fed by the Shopping Graph; structured data is how you enter it.
  3. Embed IPTC metadata at the pipeline level. Creator, credit, and license fields travel with the file. This is cheap at ingest and impossible to retrofit across a million delivered assets later.
  4. Favor original, information-dense visuals over stock. Google now generates custom images inside AI Overviews with its Nano Banana image model (2026), which means generic stock photography is being commoditized by the answer engine itself. Original charts, diagrams, and product photography carry information a model cannot synthesize and still earn citations. A chart with your data on it is a citation magnet; a stock handshake photo is landfill.
  5. Keep image performance inside the SEO budget. Fast, responsively marked-up images are documented quality signals, and they are also the difference between your asset and a competitor's when an AI surface picks one visual anchor.

One honest caveat from our side of the fence: at Inverity we build image infrastructure, so we are structurally inclined to believe images matter more every year. That is exactly why this post leans on Pew, SparkToro, Google's own announcements, and HTTP Archive rather than our opinion. Check the sources; the thesis should survive without us.

FAQ

Does alt text still matter for SEO now that AI can "see" images?

More than before, not less. Retrieval, the step that decides which images are candidates for an answer, remains text-first across Google, ChatGPT, and Perplexity: alt text, captions, and structured data are the handles. With only 55% of images carrying non-blank alt text (Web Almanac 2024), good alt text is now a competitive moat, not a checkbox.

How do I get my images cited or shown in Google AI Overviews and AI Mode?

Make them retrievable and machine-verifiable: descriptive alt text and filenames, ImageObject or Product structured data, image sitemaps, and 1200px+ source files per Google's guidance (Google Search Central). AI Mode's visual fan-out analyzes objects within images (Google, 2025), so clean, well-lit, single-subject images parse best. Original charts and diagrams outperform stock.

What structured data do I need for product images?

Product markup with image fields meeting Google's minimums: 250 x 250 px for apparel, 100 x 100 px for other products, under 64 megapixels, with 1200px+ width recommended for rich results (Google Search Central). Provide multiple aspect ratios where possible. For licensing visibility, either structured data or embedded IPTC metadata qualifies for the Licensable badge.

Do ChatGPT and Perplexity index images, and how do they choose which to show?

They select images through text signals on retrieved pages rather than crawling images independently. Perplexity retrieves and cites on nearly every answer, giving well-labeled images frequent exposure; ChatGPT often answers without retrieval and skews to high-authority domains. Citations concentrate heavily: Profound's 680M-citation analysis found only 11% of domains cited by both platforms (Profound, 2025).

Does image file size or page speed affect image rankings?

Yes, directly and by Google's own documentation, which lists fast loading and responsive markup as image quality signals (Google Search Central). Since the LCP element is an image on most pages, image weight also feeds Core Web Vitals, which affects page-level rankings that image visibility inherits. Compression strategy and search strategy are the same strategy.