Google Lens API: The Search Google Doesn't Make Public, Plus the Tools to Refine It
Google doesn't offer a public Google Lens API. See what your options are, why raw Lens results fall short, and how to turn them into exact product matches.
“Pick a category”, “write unique product descriptions”, “optimize your title tags”.
Every piece of standard ecommerce SEO advice is correct, and above a certain catalog size every piece describes work that a small team is never going to be able to handle.
Practitioners have been saying so to each other for years. How Are You Handling SEO for E-Commerce Sites with Thousands of Products? is a real thread with real answers in it, and the answers mostly amount to: prioritize the categories, canonicalize the variants, and accept that unique copy on every product page is not happening.
This guide is for that kind of ecommerce. Ten thousand to several million SKUs, vendor feeds, variants, filters that multiply URLs faster than anyone can audit them, and a team of two or three people responsible for all of it.
At that size the thing blocking your rankings is almost never what the checklists say it is. Crawl budget and faceted navigation are real problems, they are also solved problems, and fixing them gets you crawled rather than ranked. The two things that actually decide whether a large catalog ranks are both data problems.
This article covers how to actually handle this work at scale using ai.
Think about the scale question. A store with 5,000 products and ten active filter types, color, size, brand, rating, price band, material, fit, availability, sale status and shipping speed, can easily have 50,000 unique URLs. Not theoretical URLs: Crawlable ones for Google to index and rank. Botify has reported a site with under 200,000 products exposing more than 500 million pages to search bots through unconstrained filter combinations.
Then the number that decides everything below it, which no tool reports: the SEO team is two people and the catalog is 80,000 SKUs. Split up that is 40,000 product pages each. At a brisk ten minutes per page to research the product, check the attributes, write a description, set a title and meta, and hit publish, one person clears 48 pages a week. The catalog takes sixteen years.
That arithmetic is the whole article. Everything that follows is sorted by one question: does this scale, and if not, what does.
This half is generally solved, and the best public treatment of it is Go Fish Digital's ecommerce SEO checklist, written by a practitioner who has audited sites at this size. If technical SEO is the only thing you came for, read that and skip ahead. What follows is the short version and the reason it is not where large catalogs lose.
One measurement is worth taking before anything else: what share of crawl requests land on URLs you want indexed. Server logs are the only place that answer lives, because Search Console reports sampled aggregates and will not tell you a third of your crawl is going to sort-order parameters. If parameter URLs are taking more than a third, that is the first fix, and the section below is why.
Filters are where a manageable catalog becomes an unmanageable URL space. The framework that has settled as standard practice sorts every filter type into three tiers:
How few pages survive tier one is the whole discipline: a few dozen to a couple of hundred against thousands that are technically possible. Done properly it is reported to add 15% to 30% more long-tail traffic, and not because filter pages are magic. Each one is a real landing page carrying copy that is not the parent category's copy, which is the rest of this article arriving early.

As you already know SEO is not a set and forget workload. You are constantly optimizing and adjusting the SEO work to fit the changes to search. Not only that, but you have to track all of these pages in google to decide what needs adjustment.
This further increases the issues with handling SEO in large catalogs.
Search engines rank what is on the product page. That sentence is so obvious it gets skipped, and skipping it is why technical webpage fixes on large catalogs so often produce nothing. You spend a quarter on crawl budget, Googlebot starts reaching the long tail efficiently, and rankings do not move, because the pages it is now reaching efficiently have nothing on them worth ranking.
Four things decide whether a product page can compete, and we’ve covered the per PDP items before. At catalog scale the interesting question is not what they are, it is where each one comes from:
Three failure modes, and most large catalogs are running all three simultaneously on different parts of the catalog.
Which brings the conversation back to where this article opened. The practitioner threads reach that resignation because it was the correct call for as long as the only way to produce product content was a person sitting down and writing it. That constraint is the assumption worth re-examining, and it is the only one in this article that has actually changed.
At 500 SKUs a merchandiser notices that a supplier sent through a spec sheet with the wrong voltage. At 80,000 nobody has read most of your product pages, including you. The errors are not more frequent, they are less visible, and they compound in a specific way: a wrong or missing attribute removes the product from every filtered view that uses that attribute, kills it for every comparison query that mentions it, and produces structured data that describes nothing. One empty field on one page is a rounding error. The same field empty across 12,000 pages is a category you do not rank in.
Let’s talk about how Pumice.ai automates this work for large catalogs every day and actually sees SEO results at scale.

Pumice starts with the Merchandising Pipeline. The goal is to take whatever product data you have now, and turn it into complete, accurate, unique page content for every SKU. Pumice runs that as a pipeline in six stages, and the order matters because each stage constrains the next.
Research. For each product, find the live source. Universal Search locates the manufacturer or vendor page from the title, the MPN and the brand, Smart Scrape pulls the structured content off it, and PDF line sheets and supplier feeds are parsed into the same shape. The output is a verified product record rather than an assumption. This is the stage that makes everything downstream grounded rather than generated, and skipping it is what separates a content tool from a data pipeline.
Categorization to your taxonomy. Every SKU is assigned to the correct node in your site's category structure. This is the stage people underestimate, because site architecture at scale is categorization at scale: the shallow, logical hierarchy the technical section asked for only exists if every product is in the right place in it. Width has run this at 97% top-level accuracy across a 5,585-category marketplace taxonomy, on 50 million records a month, and at 92% accuracy at the deepest level of a four-level multilingual taxonomy for a wholesale marketplace whose source data arrived poorly labelled.
Attribute enrichment. Fill the fields the category requires from the verified source: material, dimensions, compatibility, capacity, certifications. Attribute completeness is what drives filter pages, structured data and comparison queries, and it is the single highest-leverage thing to fix on a large catalog because one enrichment pass improves the product page, the category page, the filter pages and the schema simultaneously.
Rules and examples configuration. Per category, you define the title pattern and its character budget, the attribute schema, the tone, the words that are never allowed, and two or three example outputs that show what good looks like. This is written once per category and applied to every SKU in it, which is what keeps 80,000 pages consistent without making them identical.
Generation. Title, description, attributes and meta per SKU, written against the verified record and the keyword roadmap together. The description is written from the specifications that were actually retrieved for this product, so two products in the same category with different specifications get substantively different pages rather than the same page with the numbers swapped.
Validation, then publish. Character limits, required attributes present, no forbidden terms, no duplicate phrasing against the rest of the catalog, schema valid. A record that fails goes back to generation rather than forward to the site. Then the batch publishes.

This distinction decides whether the output ranks. Templating arranges the fields you already have, so a missing material stays missing and every page in the category comes out structurally identical, because the template is the page. Grounded generation researches the product first and writes from what was found, so two units in the same category produce different pages and anything the pipeline could not verify is flagged rather than filled in with something plausible. A field is either true or empty, never invented.
A general-purpose model given a product name and asked for a description produces something fluent and confident. It also produces specifications inferred from the product category rather than read from your supplier, because that is what a language model does when the facts are absent. On a marketing page that is a style problem. On a product page it is a returns problem, a chargeback problem, and a low-quality signal when the description contradicts the specification table beside it. The fix is not a better prompt but retrieving the real product data before any text is generated, which is what the research stage is for.
Suppose the first problem is solved. Every SKU has complete attributes, a grounded description, a title built from real demand. You are in better shape than most large catalogs on the web. You are also not finished, because product page optimization is not a project that completes.
Four things move underneath a product page after it publishes. Competitors rebuild their pages and start covering attributes you do not. Demand shifts, seasonally and with the product cycle, so the terms that led your title in March are not the leading terms in September. Google reshuffles the result set and the page that sat fourth is now ninth without anything about it changing. And your own catalog turns over, with new SKUs arriving and discontinued ones lingering.
Every ecommerce SEO guide handles this with a line near the end about monitoring Search Console and iterating. For a 200-page site that is a reasonable instruction. For 200,000 it is not an instruction at all, because the work it describes cannot be done by the people being instructed to do it. Search Console will happily tell you that 43,000 product pages are getting impressions and no clicks. It will not tell you which of them to fix first, or what is wrong with them.
The universal fallback is to sort by revenue, optimize the top few hundred products by hand, and revisit them quarterly. It is a rational allocation of a scarce resource and it surrenders the reason large catalogs exist.
Head terms in any category are contested by every competitor with a budget, which is why they are expensive in paid and slow in organic. The long tail is the opposite: specific, high-intent, low-competition queries spread thinly across tens of thousands of products, worth little individually and a great deal in aggregate. That traffic is only reachable if the pages exist, are complete and stay current. The top 500 strategy optimizes precisely the pages where organic is hardest to win and abandons the pages where it is easiest.
The constraint is throughput, not knowledge. The team knows what to do to any given page. They cannot do it to forty thousand of them, twice a year.

The Product Optimization Playbook runs the same analysis a good SEO would run on one product page, against every product page, on a schedule. For each SKU it takes the live page and compares it to the pages currently ranking for that product's terms, then reports what is missing.
The output is a prioritized list of specific changes per SKU, not a score. A score tells you that page 41,338 is a 62. A change list tells you that page 41,338 is missing three attributes the top three results all carry, and that its title leads with a term whose volume fell by half.
This is the part that makes a two-person team effective rather than merely busy. Sorting by current traffic sends you to pages that are already working. Sorting by revenue sends you to the head, which is contested. The useful sort is revenue-weighted impression share: pages that are being shown for valuable queries and not being clicked, or shown in positions just below where clicks begin.
Those pages have already proved the demand exists and the page is eligible. They are losing on execution, which is the thing you can actually change. On most large catalogs that filter reduces 80,000 SKUs to a working set of two to three thousand, which is a quarter of real work for a small team rather than an impossible backlog.
Findings go back into the Merchandising Pipeline as a regeneration batch rather than a to-do list. The attributes that were missing get retrieved and filled, the titles that drifted get rewritten against current demand, the meta that underperformed gets rebuilt. Then it runs again, because the same forces that moved the pages the first time have not stopped.

The difference between this and the standard advice is that nothing here waits on a person deciding which page to look at next. A person decides the rules, reviews the batch and approves what ships. The search for what is wrong, across the whole catalog, runs on its own.
Category pages carry the terms with real volume and collect internal links from every product beneath them, and on most large catalogs they are an H1, a product grid and nothing else. The cause is the same throughput problem: 900 nodes need 900 pieces of distinct copy, so the top thirty get written and the rest get a templated sentence.
The fix is the same one too. Once every product in a node has complete attributes, the category copy can be derived from what is actually in it, the brands stocked, the range of sizes and materials, the price bands, and regenerated when the assortment shifts rather than going stale the moment a supplier changes. It is the categorization layer doing double duty: putting products in the right node, then describing the node from its contents.
ChatGPT, Claude and PerplexityAi are hitting the same URLs as Google, including every filter combination you left open. Two numbers are worth holding together, because most coverage quotes the first and omits the second. AI crawlers are now roughly a fifth of bot traffic, with Cloudflare measuring around 20% of verified bot traffic in mid-2026. But close to 90% of that crawling is training rather than answering, and only a low single-digit share responds to a live user query.
So the honest version of the AEO case is narrower and more durable than the usual pitch. Most AI crawling of your catalog is corpus building, and what goes into the corpus is whatever your product data says. Optimizing for search and optimizing for answers converge at catalog scale, because both reduce to the same requirement: complete, accurate, structured product data that states the specifications plainly rather than implying them in prose. If the underlying data is thin, no amount of generative engine optimization tooling fixes it.
Run them on a fixed cohort of SKUs, recorded before any changes and again after a full cycle for that category. Comparing the whole catalog to itself over time tells you about seasonality. Comparing a fixed cohort tells you about the work.
The reason generic ecommerce SEO advice fails on a large catalog is not that the advice is wrong. Unique product content, complete attributes, clean architecture and sensible crawl control are correct at every size. The advice simply assumes a number of pages a person can hold in their head, and above that number every instruction in it becomes a description of work nobody is going to do.
So sort the work by whether it scales. Technical hygiene does, because it is template and configuration work that happens once and applies everywhere, and it gets your catalog crawled. Product data quality and per-SKU optimization do not scale by hand at all, which is why they are where large catalogs lose, and why they are worth running as pipelines instead. Fix the data underneath the pages, then keep the pages current against a result set that will not hold still.
Send us a slice of your product data, an ERP export, a supplier feed or a crawl of your live PDPs, and we will tell you what share of it is duplicate vendor copy, where the attributes are missing, and what a grounded version of those pages would look like.
Sort every filter type into three tiers. Build clean indexable URLs for combinations with proven demand, validated at 50 or more monthly impressions in Search Console or through internal site search. Canonicalize near-duplicates such as sort orders to the parent category. Block infinite parameter space, price sliders and pagination crossed with sort orders, in robots.txt so it is never crawled. Layer the signals rather than relying on one, and expect the indexable set to be small: a few dozen to a couple of hundred pages.
Not by writing them, and not by templating them either. Templating arranges the fields you already have, so every page in a category comes out structurally identical and the missing fields stay missing. Grounded generation researches each product against its real manufacturer source, enriches the attributes from what is published there, then writes from that verified record against per-category rules. Different specifications produce different pages, and anything unverifiable is flagged instead of guessed at.
It writes fluent product copy, which is the smaller half of the job. Given a product name and no specifications, a general-purpose model fills the gaps with what is statistically likely for the category, which is how catalogs publish a capacity, a material or a compatibility the product does not have. On a product page that is a returns problem, not a style problem. The fix is retrieval before generation: get the real product data first, then write from it. A pipeline, not a prompt.
The economics are better at scale than anywhere else. A large catalog has thousands of low-competition queries nobody is bidding on, and the marginal cost of one more ranking page is close to zero once the pipeline exists. What has changed is that the same structured product data now feeds answer engines as well as search results. What has not changed is that thin and duplicated pages do not rank.