Programmatic SEO is building many pages from one template and one dataset, each page aimed at a search that follows the same pattern, such as "software development company in Dallas" or "Cursor vs Claude Code". It turns into a doorway network when the rows run out of real differences and the pages differ only by the word that was swapped in. That line can be measured, and we measure it on every build.
Every programmatic search has a fixed part and a changing part
Look at the searches people type for a service or a product and a pattern shows up quickly: "[service] company in [city]", "[product] alternatives", "[A] vs [B]" and "[app] and [app] integration" are four of them. The template carries the part that stays the same. A dataset carries the part that changes, one row per page.
We keep a catalog of twelve of these patterns, which we call page families: locations, industries, hiring a role, comparisons, alternatives, integrations, templates, examples, tools, directories, explainers and answers. Each family lists the facts every row has to carry before its page is allowed to exist. A row that's missing one is skipped, and it never ships as a thin page.
The families don't carry the same risk. Comparisons are low risk, because two products differ by nature and each pair brings its own facts. Locations are the riskiest pattern we know of. It's the canonical doorway shape, and it survives only on facts that are true of one metro and false of the next.
What Google calls a doorway
Google's spam policies define it in one line: "Doorway abuse is when sites or pages are created to rank for specific, similar search queries." One of the listed examples could describe most city networks on the web: "Having multiple domain names or pages targeted at specific regions or cities that funnel users to one page."
The same policy page covers scaled content abuse, "when many pages are generated for the primary purpose of manipulating search rankings and not helping users." It says this applies "no matter how it's created." We read that as a warning about the whole family. A set of near-identical pages is judged as a set, so one lazy template puts every page built from it at risk.
Our city network, and the 67% that nearly sank it
In September 2026 one of our own sites published 77 city pages. They covered 24 metros for AI development, software development and Databricks consulting, plus app development in the five metros where that phrase measured: Dallas, Chicago, Houston, Austin and New York. The searches behind them are commercial investigation. There's no map pack on "AI development company Denver", so an office in Denver isn't a ranking input.
The first build shared 67% of its sentences between two cities. That's a doorway by any reading of the policy above, so we stopped and wrote the part that couldn't be templated. Every metro page now carries:
- what the metro's economy runs on
- why that maps to the service, written separately for AI and for software because they're different arguments
- the compliance reality that bites there, such as Part 500 in New York, 21 CFR Part 11 in Boston, TISAX in Detroit, ITAR in Denver and CMMC in San Diego
- the question that market asks, answered with its own facts
- the published case study a buyer there would recognise
- working-hours overlap, which differs by zone because Arizona skips daylight saving
That took the pairwise figure from 67% to 50%. Every page also says plainly that the company's only office is in Gilbert, Arizona. A page that implies an office it doesn't have is a worse problem than a thin one.
Measuring the line on the built pages
A rule kept in someone's head drifts the first time a deadline arrives, so the check runs on the built HTML. It takes every sentence of eight words or more on a page and counts how many also appear on a sibling page in the same family. Above 80% the page fails and the build stops.
The 80% ceiling is a different measure from the pairwise 50% above, and it was calibrated on the forty original city pages, which measured between 74% and 76% by this method. The shared half of a page is shared on purpose. The description of the service, the engagement steps and the FAQs about the service itself don't change by city, and rewording them per page to look different makes worse copy and fools nobody.
Two signs that Google had noticed
The first came in mid-September. For "databricks consulting san diego" Google served our Los Angeles page, and for Miami it served Boston. It was reading the Databricks city pages as interchangeable, which told us that family needed more per-metro argument before another metro was added.
The second came from Search Console on 20 September. Of 25 Databricks city pages, 13 sat in "Discovered, currently not indexed", against 2 of 25 software pages and none of the 25 AI pages. Part of the cause was demand, because every Databricks city term we measured read zero. Part of it was our own linking.
Each city page linked to the first nine metros in a fixed list. Chicago and New York collected 28 inbound links each, while Austin and ten other cities got 4. The fix was to link every sibling from every page, or rotate the set, and never take a fixed first slice of an ordered list. Austin and Phoenix were indexed by 28 September, after the link fix and a manual crawl request.
Publish only the rows where demand measures
A keyword study on 11 September 2026 crossed 51 phrasings with 39 metros in three word orders. Of roughly 12,500 monthly searches across every city term that measured, 12,260 were custom software and app development. AI agent and Databricks city terms measured zero everywhere.
So the app development family publishes five metros and stops. Zero from an exact-match tool doesn't always mean nobody searches, since those tools round small numbers away, and some of our zero-volume city pages were still quoted in AI Overviews. But each near-duplicate page adds exposure, and a page should exist where there's demand for it.
Three questions before you build a family
We ask these in order, and the third one kills most ideas.
Is the person searching a buyer? A page that ranks for students brings traffic and no enquiries. Can the company give that buyer something useful? An integration page needs an integration that exists. Is there distinct material for every page? If two pages would differ only by the swapped word, you don't have a family. You have a doorway network, or four honest pages.
Then start small. Our engagements build five to ten pilot pages from the dataset, run them through every check, and wait for real results before adding a second pattern.
These rules, the twelve family specs and the uniqueness check now live in Heddle, the toolkit we use to set up programmatic pages on our own products and for clients.
Questions
What is programmatic SEO in one sentence?
Take one template and one dataset, and aim each row at a search with the same shape, such as "software development company in Dallas". The template is shared.
What makes each page worth reading is its row, the facts that city has and its siblings don't. That row is where all of the work is.
Is programmatic SEO against Google's guidelines?
The method isn't. Google's spam policies target pages created to rank for similar queries that aren't useful in themselves, and pages generated at scale to manipulate rankings. A programmatic page with facts true of its own row is a page. One that differs only by the swapped word is the doorway the policy describes.
How different do programmatic pages need to be?
We fail any page that shares more than 80% of its sentences of eight words or more with a sibling, measured on the built HTML. Our city pages sit between 74% and 76% on that measure, because the service description and steps are shared on purpose.
How many pages should a programmatic SEO project start with?
Five to ten. Build a pilot from the rows with the most demand and the best material, put it through every check, and wait for real results before adding rows or a second pattern.