We audited our own local directory. About 9% of its 10,632 web addresses were pages worth finding.
Local Times OKC is a directory of Oklahoma City area businesses. It has the same owner as OKC SEO, and it runs on Brilliant Directories, a hosted directory platform. On October 7, 2026 we counted every web address the platform's sitemap listed for our site: 10,632. About 956 of them, roughly 9%, were pages we would want a searcher to land on: 910 business listings and 46 pages we wrote on purpose: 37 category hubs plus the home, about, how-we-list, contact, legal and index pages.
Most of the rest came from features we never set up on purpose. There were 6,543 search-result pages covering about 935 real topics, around seven addresses per topic. There were 1,820 contact-form and write-a-review pages, 377 dead pages created by spam sign-ups, and 13 post-type pages, 12 of them empty. Our robots.txt file also blocked ChatGPT's search crawler, Perplexity, Claude and Apple. We changed what we could the same day, and the sitemap dropped to 8,813 addresses. Next, we plan to move the directory to our own build.
These are the defaults and settings we found on one site, ours. The platform has settings that change several of them, and another directory on the same platform can look very different depending on how it is set up.
Disclosure. Local Times OKC and OKC SEO have the same owner. We use Brilliant Directories as a paying customer. We have no affiliate or partner relationship with the company and receive no payment, commission or other benefit from it. Nobody paid for this study, and the platform did not review it before publishing.
- 10,632addresses in the platform's sitemap
- 956 (9%)pages worth landing on
- 6,543search-result pages for about 935 topics
- about 2.9 MBper directory page, vs 14 to 74 KB on our own builds
- 8,813addresses left after same-day fixes
Method and limits
- What we measured
- Every web address in the XML sitemap the platform generated for www.localtimesokc.media (one index file and 40 child files). For a sample of those addresses we also recorded the HTTP status, the robots meta tag, the canonical tag, how many business listings the page showed, how many words were unique to the page, and the page weight. We also read the site's robots.txt file and its platform settings.
- When
- October 7, 2026. The before counts were taken first; the after count was taken once the sitemap had been regenerated, and re-checked later the same day.
How
- Downloaded the sitemap index and all 40 child files and counted the addresses in each file. Each child file holds one kind of page, so the file names give the page types.
- Fetched up to 5 addresses from every child file (191 pages), plus a random sample of 120 of the 6,543 Oklahoma search-result addresses, one page request every 1.2 to 1.5 seconds with a user agent that named us. Read-only: nothing on the site was changed during the count.
- Counted unique words as the text blocks that appear on fewer than 30% of the sampled pages. That leaves out the menus, sign-in boxes and footer that repeat on every page.
- Measured page weight as the uncompressed size of the HTML plus the scripts, stylesheets and images the page loads. Fonts, map tiles and files that load later were left out.
- Grouped the search-result addresses by city and service, so the versions with and without /united-states/ or /oklahoma/ in the address count as one topic.
- Read the platform's settings and listing records through its API (read-only) and read its support articles for each setting.
Limits
- One site on one day. Our site had a history of spam sign-ups, cleaned up in early October 2026, and those sign-ups inflated some counts.
- The share of search pages showing one listing and the share returning Not Found come from samples, so treat them as estimates. The topic count is a full count of all 6,543 addresses.
- We counted what the sitemap lists, not what Google has indexed. Our Search Console access for this site started on October 7, 2026, so indexing numbers come in a later update.
- We did not run Lighthouse or measure Core Web Vitals. The page-weight figures are bytes and response times only.
- The weight comparison is with our own simpler sites, which have no member logins, maps or sign-up forms. Part of the gap is features those sites don't have.
- We counted all 910 listings as useful because each was an active Oklahoma City area business on the platform when we counted. About 615 of them have not yet been checked against the business's own website. Later on October 7 we hid 8 that failed that check (no working website or phone), and more may follow.
Where the 10,632 addresses came from
The platform builds its sitemap from templates. Each template turns into one or more child files, and each file lists every page that template can make. Here is every file type, grouped by what the pages actually show.
| Page type | Example address | Addresses | What the page shows | Search setting on Oct 7 |
|---|---|---|---|---|
| Business listing | /oklahoma-city/plumber/business-name | 910 | The business's own page: description, services, contact details | Indexable. Useful |
| Hub, about, contact, legal and index pages | /okc-plumbers | 46 | 37 category guides we wrote and checked, plus the home page, how we list businesses and other site pages | Indexable. Useful |
| Utility pages | /passwordreset, /sitesearch | 7 | Password reset, site search, search results, a get-matched form, a data feed and 2 redirects | Mostly indexable |
| Contact form for each listing | .../business-name/connect | 910 | A form, with 38 to 48 words of its own | Indexable |
| Write-a-review form for each listing | .../business-name/writeareview | 910 | A form, with 65 to 69 words of its own | Indexable |
| Review list for each listing | .../business-name/reviews | 910 | Empty. The whole site had 5 reviews, all spam from 2021 | Marked noindex, yet listed in the sitemap |
| Single review pages | .../reviews/38 | 4 | Spam reviews on deactivated profiles | Marked noindex, yet listed |
| Search results: Oklahoma places, or no place | /oklahoma-city/leak-detection | 6,543 | 1 to 10 listing snippets and no introduction | Indexable. Each names itself the preferred address |
| Search results: places outside Oklahoma | a city name added by a spam sign-up | 377 | Nothing. 404 Not Found | Dead |
| Post-type pages | /coupons, /jobs, /videos | 13 | 12 were empty; one had a single 2019 post | Indexable, except one |
| Membership plan page | /copy-4 | 1 | An automatic page for a plan we use internally | Indexable |
| Photo post | one product post | 1 | A single 2019 post | Indexable |
| Total | 10,632 | 956 useful (9%) |
Useful = the 910 listings plus the 46 hub, about, contact, legal and index pages. 956 / 10,632 = 9.0%. Word counts, statuses and robots settings come from the sampled pages.
Two more details matter for anyone reading their own sitemap. First, none of the search-result files had been submitted to Search Console, but that does not hide them: our listings link to them under "Related Searches" and in their breadcrumb trails. Second, on our 10-listing search pages the page-number links were empty in the HTML, so a crawler that doesn't run JavaScript could not page past the first 10. We have not checked whether the links appear once JavaScript runs.
One topic, about seven addresses
The biggest group was search results. The platform makes a page for every combination of country, state, city, category and service it knows, with and without each part of the address. For plumbers who do leak detection in Oklahoma City, these four addresses all loaded and each named itself canonical: /oklahoma-city/leak-detection, /united-states/oklahoma-city/leak-detection, /oklahoma/oklahoma-city/leak-detection and /united-states/oklahoma/oklahoma-city/leak-detection. The first two showed the same 12 businesses; the two with /oklahoma/ in the address showed 8 of those 12. Add the same four with /plumber/ in the address and this one topic has eight. Across all 935 topics the average is about seven (most have four or eight).
We grouped the 6,543 Oklahoma addresses by city and service and found about 935 real topics: 307 for Oklahoma City, 150 for Edmond, 105 for Moore and 373 with no city. Every version we sampled named itself as the canonical (the preferred address), so the site never told search engines which one to pick.
Many of those pages are also small. In our random sample of 120 search-result pages, more than a third showed a single business.
| Listings on the page | Pages in sample | Share |
|---|---|---|
| 1 | 46 | 38% |
| 2 to 3 | 21 | 18% |
| 4 to 9 | 32 | 27% |
| 10 (the most one page shows) | 21 | 18% |
| All sampled pages | 120 | 100% |
Shares are rounded, so they add to 101%. All 120 returned 200 OK, were set to index, follow, and named themselves canonical. In a separate sample of 101 Oklahoma search addresses, 2 returned 404, so roughly 2% of these pages may be dead as well.
"If you have the same content accessible under different URLs, choose the URL you prefer and include that in the sitemap instead of all URLs that lead to the same content."
1,820 form pages that could be indexed
Every listing came with two more indexable pages: a contact form and a write-a-review form. Each had a few dozen words of its own, and the rest was the site's menus and footer. The platform's own help article describes these pages well.
"These pages are generally intended to support actions rather than provide unique content and often contain little or no standalone information."
The platform has a setting for exactly this, called Noindex Review & Connect Pages. It was off on our site. When it is on, the article says, these pages carry noindex, follow, which keeps them out of search results while search engines can still follow their links.
377 dead pages from spam sign-ups
Spam accounts had signed up with addresses outside Oklahoma. Each sign-up added its city to the site's city list, 38 cities in all, and the platform then made search-result pages for those cities. All 54 we sampled returned 404 Not Found, yet the sitemap still listed 377 of them.
The 910 review-list pages had the opposite problem. They were marked noindex, which tells search engines not to show them, while the sitemap asked search engines to crawl them. Google's sitemap guide says to include only the addresses you want shown: "Include the URLs in your sitemap that you want to see in Google's search results."
Page weight: about 2.9 MB a page
Every directory page carried the full template: a sign-in box, a newsletter box, two menus and the footer, plus our own call-line widget. That came to about 500 words of repeated text. An empty Not Found page had 516 words. Here is how three directory page types compared with pages from two other directories we run on a lean static build (Astro on Cloudflare).
| Page | HTML | Total loaded | Script tags | External scripts | Stylesheets | HTML response time |
|---|---|---|---|---|---|---|
| Directory search-result page | 376 to 411 KB | 2,895 KB | 53 | 10 | 12 | 0.38 to 0.53 s |
| Directory business listing | 436 to 442 KB | 2,947 KB | 67 to 70 | 12 to 13 | 12 | 0.49 to 0.58 s |
| Directory hub page | 288 to 290 KB | 2,791 KB | 45 | 8 | 12 | 0.32 to 0.33 s |
| Our Christmas market directory, market and state pages (Astro) | 9 to 17 KB | 14 to 23 KB | 2 to 4 | 0 | 1 | 0.18 to 0.23 s |
| Our hotel pool directory, city page (Astro) | 68 KB | 74 KB | 3 | 0 | 1 | 0.23 to 0.33 s |
Total loaded = uncompressed HTML plus the scripts, stylesheets and images the page requests; fonts, map tiles and later-loading files are left out. 1 KB = 1,024 bytes. Each directory page also carried about 150 to 280 KB of inline JavaScript and about 100 KB of inline CSS.
Speed is not the whole story, and our lean pages do less: no logins, no maps, no sign-up forms. But for a visitor on a phone who wants a plumber's number, most of those 2.9 MB are features they never use.
robots.txt shut out AI search crawlers
The robots.txt file our site was serving named a list of crawlers (Google's, Yahoo's, Microsoft's MSN bots and Twitter's) and let them in. Its last rule, User-agent: * with Disallow: /, blocked every crawler it did not name. It also had no Sitemap line. We have not checked whether other sites on the platform start with the same file, so check yours.
| Crawler | Run by | Before | After |
|---|---|---|---|
| Googlebot | Google Search | Allowed | Allowed |
| Msnbot / Bingbot | Microsoft Bing | Msnbot allowed; Bingbot not named in the served file | Both named and allowed |
| OAI-SearchBot | OpenAI (ChatGPT search) | Blocked | Allowed |
| GPTBot | OpenAI (content that may be used for model training) | Blocked | Allowed (owner's choice) |
| PerplexityBot | Perplexity | Blocked | Allowed |
| ClaudeBot | Anthropic (content that may be used for model training) | Blocked | Allowed |
| Claude-SearchBot | Anthropic (Claude's search results) | Blocked | Allowed |
| Applebot | Apple | Blocked | Allowed |
| DuckDuckBot | DuckDuckGo | Blocked | Allowed |
| Sitemap line | Tells crawlers where the sitemap is | Missing | Added |
What OpenAI's and Anthropic's crawlers do comes from each company's own crawler page (see Sources). The admin copy of the file also had a Bingbot group that the live file did not show. Only /api/ is blocked now, for every crawler.
"Sites that are opted out of OAI-SearchBot will not be shown in ChatGPT search answers, though can still appear as navigational links."
Robots.txt controls crawling, not indexing. Google's guide is direct about it: robots.txt "is not a mechanism for keeping a web page out of Google". To keep a page out of search, use noindex instead. That is why we fixed the thin pages with noindex rather than blocking them.
What we changed the same day
The owner approved a list of quick fixes before any were made. These five were made, each logged with a way to undo it. Business listings were not touched.
| Change | Result |
|---|---|
| Opened robots.txt to all crawlers except /api/ and added the Sitemap line | Live file checked byte for byte; every crawler in the table above may fetch our pages |
| Turned on Noindex Review & Connect Pages | Contact and review-form pages now carry noindex, follow; their 1,820 addresses left the sitemap |
| Marked 11 empty post-type pages noindex | /products kept, because it has a real post |
| Removed the post-type sitemap file from Search Console | Two sitemap files remain submitted: listings and our own pages |
| Regenerated the sitemap | 10,632 addresses in 40 files became 8,813 in 38 files (re-checked the same day, before the 8 listings mentioned under Limits were hidden; the sitemap has not been regenerated since) |
Still open: the 914 noindex review pages are still in the sitemap, the 6,543 search-result pages are still there, and removing the 38 spam cities needs the platform's support team. We have drafted that request, along with a question about whether whole search-result patterns can be set to noindex.
What we are doing next
We plan to move Local Times OKC to our own build, the same Astro and Cloudflare setup our other directories run on. The rules we want are simple. One address per real topic. A hub page only where at least 3 verified businesses exist. Nothing in the sitemap that we don't want found. Every old search-result address redirected to the matching hub or marked as gone. Every listing we keep stays at its current address.
We will add the results to this page: first the Search Console indexing numbers for the current site, then before-and-after numbers once the move has settled.
What this means for local businesses
Most business owners don't run a directory, but most appear on several, and many run their own site on a platform that can generate pages automatically. Here is what we would take from this audit.
- On a directory, your listing page is the page that counts. When you search for your business, check that the result is your listing, not a contact form or a search-result page with your name on it.
- A sitemap many times larger than the number of real pages is a sign of automatic pages. One of the reasons Google gives for choosing a preferred address is "To avoid spending crawling time on duplicate pages." It doesn't make a directory bad, but it can bury the pages that matter.
- If you want to show up in ChatGPT, Perplexity or Claude answers, read your robots.txt. One catch-all Disallow line can shut those crawlers out without anyone noticing.
- Before paying for a fix, look in your platform's settings. Our worst problem pages had a switch.
Check your own directory or website
You don't need any tools beyond a browser and, for step 3, a free Search Console account.
Count your sitemap
Open yoursite.com/robots.txt and find the Sitemap line, or try yoursite.com/sitemap.xml. Most sitemaps are an index of smaller files; open a few and look at how many addresses each holds. Compare the total with the number of pages you would actually want someone to land on. Ours was 10,632 against 956.
Open five random addresses from it
If you land on forms, empty pages, search results with one item, or Not Found pages, those addresses should come out of the sitemap or carry noindex.
Read Search Console's Page indexing report
Look for "Duplicate without user-selected canonical" and "Crawled - currently not indexed". Google defines the first as a page that is a duplicate of another, "although it doesn't indicate a preferred canonical page". Lots of those means several addresses for the same thing.
Look for the same page at several addresses
Try your city and service with and without the state or country in the address. If each version loads and names itself canonical, pick one, keep it in the sitemap, and point the others at it.
Read your robots.txt
Look for User-agent: * followed by Disallow: /. Then check that OAI-SearchBot (ChatGPT search), PerplexityBot and Claude-SearchBot (Claude's search) are not blocked, and that the file has a Sitemap line.
Check your platform's settings before anything else
Search the settings for noindex, sitemap and robots. Ours had a switch that set 1,820 thin form pages to noindex in one step.
Want a second pair of eyes? We run this same check on directories and business websites, and we start with the free settings before anything you would pay for.
Book a free 30-minute reviewData and sources
Download the data: Local Times OKC sitemap audit, October 7, 2026 (CSV). Count of every address in the Brilliant Directories sitemap of www.localtimesokc.media on October 7, 2026, by page type, with status, robots setting and whether we judged the page useful. Aggregate counts only; no business or personal data.
Data license: CC BY 4.0. You may reuse the data with credit to OKC SEO and a link to this page. The license covers the downloadable data files; the article text is not licensed for reuse.
- Brilliant Directories support: Noindex Review & Connect Pages What the setting does and why the platform offers it.
- Brilliant Directories support: Sitemap Generator How the platform builds its sitemap; it notes that "Not all pages in a sitemap have value to Google."
- Google Search Central: Build and submit a sitemap (last updated July 8, 2026)
- Google Search Central: How to specify a canonical URL (last updated July 10, 2026)
- Google Search Central: Introduction to robots.txt (last updated December 10, 2025)
- Google Search Console Help: Page indexing report
- OpenAI: Overview of OpenAI Crawlers
- Anthropic: Does Anthropic crawl data from the web, and how can site owners block the crawler?
How this was made
The audit was run on October 7, 2026 by AI agents working for Chris S Riddle, who runs OKC SEO and Local Times OKC. The agents downloaded the sitemap, fetched the sample pages, counted them with scripts and wrote up the results. Chris approved the list of changes to the live directory before any were made.
This page was drafted with AI assistance from that write-up. Before publishing, its numbers were checked against the raw audit data by a separate AI review agent, then a second agent re-checked the page, its metadata and its sources, as described on our How we research page. We don't list AI as an author.
Reviewed by Chris S Riddle, owner of OKC SEO and Local Times OKC, on October 7, 2026.
Spotted a mistake? Email chris@okcseo.media and we will check it against the data and correct the page.
Updates
- : Published with the October 7 counts and the same-day changes.
- : Later the same day we hid 8 listings that failed a website and phone check. The counts on this page were taken before that.
Planned
- Around October 14, 2026: First Search Console numbers for the directory: pages indexed against pages submitted, and the main reasons pages were left out.
- Around October 28, 2026: Second Search Console pull, to see whether the noindexed form pages are dropping out.
- After the move to our own build has settled: Before-and-after numbers: sitemap size, pages indexed, page weight and search traffic.