Marketing

SEO for Large Ecommerce Sites: The 2026 Guide to Ranking at Catalog Scale

Matt Payne
·
September 26, 2026

“Pick a category”, “write unique product descriptions”, “optimize your title tags”.

Every piece of standard ecommerce SEO advice is correct, and above a certain catalog size every piece describes work that a small team is never going to be able to handle.

Practitioners have been saying so to each other for years. How Are You Handling SEO for E-Commerce Sites with Thousands of Products? is a real thread with real answers in it, and the answers mostly amount to: prioritize the categories, canonicalize the variants, and accept that unique copy on every product page is not happening.

This guide is for that kind of ecommerce. Ten thousand to several million SKUs, vendor feeds, variants, filters that multiply URLs faster than anyone can audit them, and a team of two or three people responsible for all of it.

At that size the thing blocking your rankings is almost never what the checklists say it is. Crawl budget and faceted navigation are real problems, they are also solved problems, and fixing them gets you crawled rather than ranked. The two things that actually decide whether a large catalog ranks are both data problems.

  • Your product data is the foundation, and it gets worse as the catalog grows. Search engines rank what is on the page. On a large catalog most of what is on the page came out of a vendor feed nobody has read, with attributes missing on the long tail and descriptions that are byte-identical to every other retailer selling the same SKU.
  • Nobody can manually optimize a catalog this size, and optimization is not a one-time job. Competitors relaunch pages, demand shifts, Google reshuffles, your catalog turns over. Every guide ends at monitor Search Console. That is a task for two hundred pages and a fiction for two hundred thousand.

‍

This article covers how to actually handle this work at scale using ai.

Key Takeaways

  • SEO for large ecommerce sites is two data problems, not one technical problem. Product data quality is the foundation, and optimization throughput is the constraint.
  • Technical hygiene at scale is table stakes. It is well documented, mostly a one-time project, and it gets your pages crawled rather than ranked.
  • Faceted navigation is the main technical risk: 5,000 products with ten filter types can generate more than 3.6 million crawlable URLs.
  • Vendor feed copy is duplicate content by definition, because every retailer carrying that SKU publishes the same paragraph. Templated titles are thin content by construction.
  • Grounded generation is not templating. Researching each product against its real source and writing from what is there produces unique pages; filling a template with feed fields does not.
  • Site architecture at scale is a categorization problem. Category, sub-category and product only works when every SKU sits in the right node, which on a large catalog is a machine learning problem with an accuracy number attached.
  • Per-SKU optimization has to run as a loop against the live results, prioritized by revenue-weighted impressions so a small team works the right couple of thousand pages first.

Why SEO for Large Ecommerce Sites Is a Different Problem

  
    ✂️ The short answer    

What makes SEO for large ecommerce sites different?

    

On a catalog of tens of thousands of SKUs or more, the constraint stops being knowledge and becomes throughput. The advice is the same advice every ecommerce site gets, but three things break at scale: faceted navigation multiplies a few thousand products into millions of crawlable URLs, product data arrives from vendor feeds that are duplicated across every retailer selling the same item and incomplete on the long tail, and no team can manually write, audit or re-optimize hundreds of thousands of product pages against a search result that changes every week. Technical hygiene is the documented half and it gets pages crawled. Product data quality and per-SKU optimization at scale are the half that decides what ranks.

  

Think about the scale question. A store with 5,000 products and ten active filter types, color, size, brand, rating, price band, material, fit, availability, sale status and shipping speed, can easily have 50,000 unique URLs. Not theoretical URLs: Crawlable ones for Google to index and rank. Botify has reported a site with under 200,000 products exposing more than 500 million pages to search bots through unconstrained filter combinations.

Then the number that decides everything below it, which no tool reports: the SEO team is two people and the catalog is 80,000 SKUs. Split up that is 40,000 product pages each. At a brisk ten minutes per page to research the product, check the attributes, write a description, set a title and meta, and hit publish, one person clears 48 pages a week. The catalog takes sixteen years.

That arithmetic is the whole article. Everything that follows is sorted by one question: does this scale, and if not, what does.

Table Stakes: Technical Hygiene at Catalog Scale

This half is generally solved, and the best public treatment of it is Go Fish Digital's ecommerce SEO checklist, written by a practitioner who has audited sites at this size. If technical SEO is the only thing you came for, read that and skip ahead. What follows is the short version and the reason it is not where large catalogs lose.

One measurement is worth taking before anything else: what share of crawl requests land on URLs you want indexed. Server logs are the only place that answer lives, because Search Console reports sampled aggregates and will not tell you a third of your crawl is going to sort-order parameters. If parameter URLs are taking more than a third, that is the first fix, and the section below is why.

Faceted Navigation: the Three-Tier Rule

Filters are where a manageable catalog becomes an unmanageable URL space. The framework that has settled as standard practice sorts every filter type into three tiers:

  • Index and build. Filter combinations that match real search demand, such as women's running shoes under $100 or size 12 wide-fit boots. These get clean URLs, unique titles, unique meta and their own introductory copy. Validate demand before you build: 50 or more monthly impressions in Search Console, a competitor keyword gap, or high-frequency internal site search with no landing page.
  • Canonicalize to parent. Near-duplicates with no distinct demand. Sort orders, three simultaneous size filters. Still crawlable so internal links flow, canonical pointing home.
  • Block at the crawl level. Parameter combinations that create effectively infinite URL space with zero search value. Price sliders generating a URL per dollar, pagination crossed with sort orders. These should never be crawled at all.

How few pages survive tier one is the whole discipline: a few dozen to a couple of hundred against thousands that are technically possible. Done properly it is reported to add 15% to 30% more long-tail traffic, and not because filter pages are magic. Each one is a real landing page carrying copy that is not the parent category's copy, which is the rest of this article arriving early.

‍

   

Sort every filter type once

  

The three-tier faceted navigation rule

  

Most large catalogs end up indexing 20 to 200 filter pages out of thousands that are technically possible. The discipline is in how few survive tier one.

                                                                                                                                                                                                     
TierWhat goes hereSignal to useCrawl impact
1. Index and buildFilter combinations with real search demand: "women's running shoes under $100", "size 12 wide-fit boots"Clean URL path, unique title, unique meta, its own intro copy, listed in the XML sitemapEfficient. Targeted crawl paths to pages that earn traffic
2. Canonicalize to parentNear-duplicates with no distinct demand: sort orders, three simultaneous size filtersCanonical tag to the parent category, or noindex where the page must stay crawlable for internal linksStill crawled, so budget is partly spent. Acceptable for a bounded set
3. Block at the crawl levelEffectively infinite URL space with zero search value: price sliders generating a URL per dollar, pagination crossed with sort ordersrobots.txt disallow on the parameter pattern. Test the rule before deployingEliminated. Never crawled at all
   

Never rely on one signal. Layer canonicals on tier two and block tier three at the crawl level, while building real URL paths for tier one. Validate tier-one demand at 50+ monthly impressions in Search Console, a competitor keyword gap, or high-frequency internal site search with no landing page.

 
Diagram showing how faceted navigation expands 5,000 products into 3.6 million URLs and how the three-tier framework reduces them
Filters multiply a mid-sized catalog into millions of crawlable URLs. The three-tier rule decides which handful survive to be indexed.

None of this accounts for ongoing SEO work

As you already know SEO is not a set and forget workload. You are constantly optimizing and adjusting the SEO work to fit the changes to search. Not only that, but you have to track all of these pages in google to decide what needs adjustment.

This further increases the issues with handling SEO in large catalogs. 

Problem 1: Product Data Quality Is the Foundation, and It Decays as the Catalog Grows

Search engines rank what is on the product page. That sentence is so obvious it gets skipped, and skipping it is why technical webpage fixes on large catalogs so often produce nothing. You spend a quarter on crawl budget, Googlebot starts reaching the long tail efficiently, and rankings do not move, because the pages it is now reaching efficiently have nothing on them worth ranking.

What a Product Page Needs Before Any of This Matters

Four things decide whether a product page can compete, and we’ve covered the per PDP items before. At catalog scale the interesting question is not what they are, it is where each one comes from:

  • A title built from real demand. Not the vendor's internal part description. The words buyers actually type for this product, ordered so the highest-demand term leads. That requires keyword data per product category, not per site.
  • Complete, accurate attributes. Material, dimensions, compatibility, capacity, finish. These drive filter pages, structured data, comparison queries and an increasing share of what AI answers quote. They are also the first thing to go missing on the long tail.
  • A description grounded in the actual product. Written from what the manufacturer published about this item, not from what a model assumes about the category.
  • Meta engineered for the result, not the page. The title tag and description are advertising copy for a SERP position, competing against nine other listings, and they follow different rules from the on-page title.

‍

   

The foundation, field by field

  

What a product page needs to rank, and what a vendor feed actually gives you

  

On a large catalog, most of what is on the page arrived in a feed nobody has read. This is the gap every technical fix is sitting on top of.

                                                                                                                                                                                                                                                                                             
FieldWhat it needs to beWhat the feed gives youWhat it costs you
TitleBuilt from real keyword demand, highest-demand term firstThe vendor's internal part description, or a template reading Brand + Name + CategoryUnique as a string, identical as a page. Ranks for the product name and nothing else
AttributesComplete and accurate against the category schema: material, dimensions, compatibility, capacityPopulated on the head, patchy or empty across the long tail, occasionally wrongDrops out of every filtered view using that attribute, every comparison query, and the schema describes nothing
DescriptionGrounded in what the manufacturer actually published about this itemThe manufacturer's paragraph, byte-identical to every other retailer carrying the SKUDuplicate content by definition. Google picks one retailer to show and it is rarely the smallest domain
Meta title and descriptionAdvertising copy for a search result, competing against nine other listingsAuto-filled from the H1, or truncated, or absent on the long tailRankings you already earned convert into clicks you do not get
Category assignmentThe correct leaf node in a shallow, logical hierarchyWhatever the person who loaded the file chose, or the vendor's own taxonomy mapped roughly acrossWrong node means wrong internal links, wrong breadcrumbs, and crawl depth nobody planned
   

At 500 SKUs a merchandiser notices these. At 80,000 nobody has read most of your product pages. The errors are not more frequent, they are less visible.

 

What Large Catalogs Actually Have Instead

Three failure modes, and most large catalogs are running all three simultaneously on different parts of the catalog.

  • Generic vendor feed copy. You imported the manufacturer's description. So did every other retailer carrying that SKU. The paragraph on your page is byte-identical to the paragraph on eleven competitors' pages, and Google has to choose one of you to show. It will not be the smallest domain. This is not a content-quality problem you can fix by editing; the text is duplicated no matter how good it is.
  • Templated titles and meta, which are thin content by construction. The standard advice for large sites is to generate title tags programmatically, and the standard result is 80,000 pages reading Brand + Product Name + Category + Buy Online. Every one is unique as a string and identical as a page. Google's quality systems have been reading that pattern for a decade.
  • Hand-written copy for the top 500 SKUs and nothing for the rest. When you have manual constraints, you focus all your time on the most important SKUs, and the others get generic or sparse copy that never ranks. 

Which brings the conversation back to where this article opened. The practitioner threads reach that resignation because it was the correct call for as long as the only way to produce product content was a person sitting down and writing it. That constraint is the assumption worth re-examining, and it is the only one in this article that has actually changed.

Why Bad Product Data Gets Harder to See as the Catalog Grows

At 500 SKUs a merchandiser notices that a supplier sent through a spec sheet with the wrong voltage. At 80,000 nobody has read most of your product pages, including you. The errors are not more frequent, they are less visible, and they compound in a specific way: a wrong or missing attribute removes the product from every filtered view that uses that attribute, kills it for every comparison query that mentions it, and produces structured data that describes nothing. One empty field on one page is a rounding error. The same field empty across 12,000 pages is a category you do not rank in.

‍

Let’s talk about how Pumice.ai automates this work for large catalogs every day and actually sees SEO results at scale.

Method 1: Rebuilding Catalog Data With the Merchandising Pipeline

Pumice starts with the Merchandising Pipeline. The goal is to take whatever product data you have now, and turn it into complete, accurate, unique page content for every SKU. Pumice runs that as a pipeline in six stages, and the order matters because each stage constrains the next.

Research. For each product, find the live source. Universal Search locates the manufacturer or vendor page from the title, the MPN and the brand, Smart Scrape pulls the structured content off it, and PDF line sheets and supplier feeds are parsed into the same shape. The output is a verified product record rather than an assumption. This is the stage that makes everything downstream grounded rather than generated, and skipping it is what separates a content tool from a data pipeline.

Categorization to your taxonomy. Every SKU is assigned to the correct node in your site's category structure. This is the stage people underestimate, because site architecture at scale is categorization at scale: the shallow, logical hierarchy the technical section asked for only exists if every product is in the right place in it. Width has run this at 97% top-level accuracy across a 5,585-category marketplace taxonomy, on 50 million records a month, and at 92% accuracy at the deepest level of a four-level multilingual taxonomy for a wholesale marketplace whose source data arrived poorly labelled.

Attribute enrichment. Fill the fields the category requires from the verified source: material, dimensions, compatibility, capacity, certifications. Attribute completeness is what drives filter pages, structured data and comparison queries, and it is the single highest-leverage thing to fix on a large catalog because one enrichment pass improves the product page, the category page, the filter pages and the schema simultaneously.

Rules and examples configuration. Per category, you define the title pattern and its character budget, the attribute schema, the tone, the words that are never allowed, and two or three example outputs that show what good looks like. This is written once per category and applied to every SKU in it, which is what keeps 80,000 pages consistent without making them identical.

Generation. Title, description, attributes and meta per SKU, written against the verified record and the keyword roadmap together. The description is written from the specifications that were actually retrieved for this product, so two products in the same category with different specifications get substantively different pages rather than the same page with the numbers swapped.

Validation, then publish. Character limits, required attributes present, no forbidden terms, no duplicate phrasing against the rest of the catalog, schema valid. A record that fails goes back to generation rather than forward to the site. Then the batch publishes.

  Alt text: Diagram of the Pumice Merchandising Pipeline turning a vendor feed into unique grounded product pages at catalog scale
Research happens first and generation happens fifth. A record that fails validation goes back to generation rather than forward to the site.

‍

Grounded Generation Is Not Templating

This distinction decides whether the output ranks. Templating arranges the fields you already have, so a missing material stays missing and every page in the category comes out structurally identical, because the template is the page. Grounded generation researches the product first and writes from what was found, so two units in the same category produce different pages and anything the pipeline could not verify is flagged rather than filled in with something plausible. A field is either true or empty, never invented.

Why Not Just Run the Catalog Through ChatGPT

A general-purpose model given a product name and asked for a description produces something fluent and confident. It also produces specifications inferred from the product category rather than read from your supplier, because that is what a language model does when the facts are absent. On a marketing page that is a style problem. On a product page it is a returns problem, a chargeback problem, and a low-quality signal when the description contradicts the specification table beside it. The fix is not a better prompt but retrieving the real product data before any text is generated, which is what the research stage is for.

Problem 2: No Team Can Manually Optimize a Catalog This Size

Suppose the first problem is solved. Every SKU has complete attributes, a grounded description, a title built from real demand. You are in better shape than most large catalogs on the web. You are also not finished, because product page optimization is not a project that completes.

Optimization Is a Loop, and Every Guide Ends Before It

Four things move underneath a product page after it publishes. Competitors rebuild their pages and start covering attributes you do not. Demand shifts, seasonally and with the product cycle, so the terms that led your title in March are not the leading terms in September. Google reshuffles the result set and the page that sat fourth is now ninth without anything about it changing. And your own catalog turns over, with new SKUs arriving and discontinued ones lingering.

Every ecommerce SEO guide handles this with a line near the end about monitoring Search Console and iterating. For a 200-page site that is a reasonable instruction. For 200,000 it is not an instruction at all, because the work it describes cannot be done by the people being instructed to do it. Search Console will happily tell you that 43,000 product pages are getting impressions and no clicks. It will not tell you which of them to fix first, or what is wrong with them.

What Teams Do Instead, and Why It Concedes the Money

The universal fallback is to sort by revenue, optimize the top few hundred products by hand, and revisit them quarterly. It is a rational allocation of a scarce resource and it surrenders the reason large catalogs exist.

Head terms in any category are contested by every competitor with a budget, which is why they are expensive in paid and slow in organic. The long tail is the opposite: specific, high-intent, low-competition queries spread thinly across tens of thousands of products, worth little individually and a great deal in aggregate. That traffic is only reachable if the pages exist, are complete and stay current. The top 500 strategy optimizes precisely the pages where organic is hardest to win and abandons the pages where it is easiest.

The constraint is throughput, not knowledge. The team knows what to do to any given page. They cannot do it to forty thousand of them, twice a year.

Method 2: Running Optimization as a Pipeline With the Product Optimization Playbook

  Alt text: Diagram of the per-SKU optimization loop from live page crawl through gap analysis and prioritization to regeneration
The loop runs against a result set that will not hold still. A person sets the rules and approves the batch; the search for what is wrong runs on its own.

The Product Optimization Playbook runs the same analysis a good SEO would run on one product page, against every product page, on a schedule. For each SKU it takes the live page and compares it to the pages currently ranking for that product's terms, then reports what is missing.

  • Attribute gaps. Fields the top-ranking pages populate that your page leaves blank. This is the most common finding and the most mechanical to fix, because the missing values are usually retrievable from the source the Merchandising Pipeline already identified.
  • Title keyword drift. Where the terms buyers use have moved away from the terms your title leads with. Seasonal on some categories, permanent on others, invisible without per-product comparison.
  • Meta underperformance. Pages holding a decent position with a click-through rate well below what that position normally returns. This is the cheapest win on any large catalog because the ranking is already earned and only the copy is losing the click.
  • Category copy gaps. Nodes where the ranking pages carry substantive category content and yours carries a heading and a product grid.

The output is a prioritized list of specific changes per SKU, not a score. A score tells you that page 41,338 is a 62. A change list tells you that page 41,338 is missing three attributes the top three results all carry, and that its title leads with a term whose volume fell by half.

Prioritize by Revenue-Weighted Impressions, Not by Traffic

This is the part that makes a two-person team effective rather than merely busy. Sorting by current traffic sends you to pages that are already working. Sorting by revenue sends you to the head, which is contested. The useful sort is revenue-weighted impression share: pages that are being shown for valuable queries and not being clicked, or shown in positions just below where clicks begin.

Those pages have already proved the demand exists and the page is eligible. They are losing on execution, which is the thing you can actually change. On most large catalogs that filter reduces 80,000 SKUs to a working set of two to three thousand, which is a quarter of real work for a small team rather than an impossible backlog.

Then Close the Loop

Findings go back into the Merchandising Pipeline as a regeneration batch rather than a to-do list. The attributes that were missing get retrieved and filled, the titles that drifted get rewritten against current demand, the meta that underperformed gets rebuilt. Then it runs again, because the same forces that moved the pages the first time have not stopped.

Snippet of a PDP optimization playbook focused on optimizing the description of a product against competitor analysis. 

  

The difference between this and the standard advice is that nothing here waits on a person deciding which page to look at next. A person decides the rules, reviews the batch and approves what ships. The search for what is wrong, across the whole catalog, runs on its own.

Category Pages at Scale

Category pages carry the terms with real volume and collect internal links from every product beneath them, and on most large catalogs they are an H1, a product grid and nothing else. The cause is the same throughput problem: 900 nodes need 900 pieces of distinct copy, so the top thirty get written and the rest get a templated sentence.

The fix is the same one too. Once every product in a node has complete attributes, the category copy can be derived from what is actually in it, the brands stocked, the range of sizes and materials, the price bands, and regenerated when the assortment shifts rather than going stale the moment a supplier changes. It is the categorization layer doing double duty: putting products in the right node, then describing the node from its contents.

SEO and AEO at Catalog Scale

ChatGPT, Claude and PerplexityAi are hitting the same URLs as Google, including every filter combination you left open. Two numbers are worth holding together, because most coverage quotes the first and omits the second. AI crawlers are now roughly a fifth of bot traffic, with Cloudflare measuring around 20% of verified bot traffic in mid-2026. But close to 90% of that crawling is training rather than answering, and only a low single-digit share responds to a live user query.

So the honest version of the AEO case is narrower and more durable than the usual pitch. Most AI crawling of your catalog is corpus building, and what goes into the corpus is whatever your product data says. Optimizing for search and optimizing for answers converge at catalog scale, because both reduce to the same requirement: complete, accurate, structured product data that states the specifications plainly rather than implying them in prose. If the underlying data is thin, no amount of generative engine optimization tooling fixes it.

‍

What to Measure on a Catalog This Size

  • Indexed versus submitted, by page type. Split XML sitemaps by category, product and filter page so Search Console reports each separately. A catalog with 60% of products indexed has a different problem from one at 95% indexed with no impressions.
  • Revenue-weighted keyword coverage. What share of the queries driving revenue in your categories have a page of yours appearing at all. The long-tail measurement, and the one that moves when catalog content goes from templated to grounded.
  • Attribute completeness against the category schema. A leading indicator rather than an outcome, and the only one here you can move directly. It predicts the other two.

Run them on a fixed cohort of SKUs, recorded before any changes and again after a full cycle for that category. Comparing the whole catalog to itself over time tells you about seasonality. Comparing a fixed cohort tells you about the work.

Conclusion

The reason generic ecommerce SEO advice fails on a large catalog is not that the advice is wrong. Unique product content, complete attributes, clean architecture and sensible crawl control are correct at every size. The advice simply assumes a number of pages a person can hold in their head, and above that number every instruction in it becomes a description of work nobody is going to do.

So sort the work by whether it scales. Technical hygiene does, because it is template and configuration work that happens once and applies everywhere, and it gets your catalog crawled. Product data quality and per-SKU optimization do not scale by hand at all, which is why they are where large catalogs lose, and why they are worth running as pipelines instead. Fix the data underneath the pages, then keep the pages current against a result set that will not hold still.

How Much of Your Catalog Has Nobody Read?

Send us a slice of your product data, an ERP export, a supplier feed or a crawl of your live PDPs, and we will tell you what share of it is duplicate vendor copy, where the attributes are missing, and what a grounded version of those pages would look like.

  
     Catalog audit    

How much of your catalog has nobody read?

    

Send us a slice of your product data, an ERP export, a supplier feed or a crawl of your live product pages. We will tell you what it actually contains.

                                                                                                   
What share of your descriptions are duplicate vendor copy published by every other retailer
Where attributes are missing against the schema each category actually needs
What a grounded version of those pages would look like, on your own SKUs
     Send us your catalog data →   

Frequently Asked Questions

How do I handle faceted navigation on a large ecommerce site?

Sort every filter type into three tiers. Build clean indexable URLs for combinations with proven demand, validated at 50 or more monthly impressions in Search Console or through internal site search. Canonicalize near-duplicates such as sort orders to the parent category. Block infinite parameter space, price sliders and pagination crossed with sort orders, in robots.txt so it is never crawled. Layer the signals rather than relying on one, and expect the indexable set to be small: a few dozen to a couple of hundred pages.

How do I write unique product descriptions for thousands of products?

Not by writing them, and not by templating them either. Templating arranges the fields you already have, so every page in a category comes out structurally identical and the missing fields stay missing. Grounded generation researches each product against its real manufacturer source, enriches the attributes from what is published there, then writes from that verified record against per-category rules. Different specifications produce different pages, and anything unverifiable is flagged instead of guessed at.

Can ChatGPT do SEO for a large catalog?

It writes fluent product copy, which is the smaller half of the job. Given a product name and no specifications, a general-purpose model fills the gaps with what is statistically likely for the category, which is how catalogs publish a capacity, a material or a compatibility the product does not have. On a product page that is a returns problem, not a style problem. The fix is retrieval before generation: get the real product data first, then write from it. A pipeline, not a prompt.

Is SEO still worth it for large ecommerce sites in 2026?

The economics are better at scale than anywhere else. A large catalog has thousands of low-competition queries nobody is bidding on, and the marginal cost of one more ranking page is close to zero once the pipeline exists. What has changed is that the same structured product data now feeds answer engines as well as search results. What has not changed is that thin and duplicated pages do not rank.

‍