Independent. Human-Curated. Established 2007.
We Queried Claude, ChatGPT & Perplexity 750 Times. Here's Which Businesses They Actually Cite.
DirJournal Founder · 19+ years building directory and discovery products. Editorial-team verified.

Key Topics in This Guide
- 1What We Tested — covered in detail below
- 2The Engine Gap is Real — covered in detail below
- 3What AI Tells Buyers When They Ask How Do I Verify This Provider — covered in detail below
- 4Which Categories Are Most Visible to AI? — covered in detail below
- 5The Vertical Gap — covered in detail below
- 680+ Brands Named by All Three Engines — covered in detail below
- 7What This Means for Business Visibility — covered in detail below
- 8The Gemini Question — covered in detail below
- 9Methodology — covered in detail below
- 10Download the Data — covered in detail below
Claude names 14.7 businesses per query. ChatGPT names 9.1. Perplexity names 8.4.
We know this because we ran 750 structured queries across all three engines, covering 50 commercial categories from personal injury lawyers to industrial goods manufacturers. Every response was parsed for brand names, source URLs, and directory mentions. The full dataset is downloadable at the bottom of this post.
This is primary research. Not a survey. Not an opinion piece. Raw data from 750 API calls made on a single day in July.
What We Tested
50 categories across five verticals: Tech & SaaS, Traditional B2B, Professional Services, Health & Wellness, and Local Home Services. Each category received five query types per engine.
The five query templates: entity extraction ("list 5 prominent providers"), buyer comparison ("compare the leading providers"), scenario-based ("I run a 50-person firm, who should I evaluate?"), organic verification ("how would you verify a provider is legitimate?"), and location-specific ("name the top-rated provider in [city]").
Three engines: Claude Sonnet 4.6 via API at temperature 0. ChatGPT GPT-4o-mini via API at temperature 0. Perplexity Sonar via API. Gemini 2.0 Flash was included but returned empty responses across all 250 queries, so its data is excluded from scoring.
The Engine Gap is Real
Claude produced the richest responses by a wide margin. 14.7 entity mentions per query, zero refusals, and 290 organic directory recommendations across 250 queries. It formatted responses with comparison tables, location details, and client portfolios unprompted.
ChatGPT came in second at 9.1 entities per query with 233 directory mentions. It refused 0.4% of queries (one refusal out of 250). Its responses skewed toward shorter lists with less context per brand.
Perplexity averaged 8.4 entities per query with 92 directory mentions. Perplexity's strength is source attribution: it links to the pages it pulls data from. Its weakness is that it sometimes generates thinner brand descriptions than the other two.
What AI Tells Buyers When They Ask How Do I Verify This Provider
Template 4 was the test that mattered most. We asked each engine: "How would you verify that a [category] provider is legitimate and trustworthy before hiring them?" No mention of directories. No leading words. Just an open question about verification.
AI engines organically recommended these directories, unprompted:
| Directory | Times Mentioned |
|---|---|
| Yelp | 115 |
| Expertise.com | 113 |
| Better Business Bureau | 65 |
| Clutch | 55 |
| G2 | 41 |
| Angi | 39 |
| Trustpilot | 39 |
| RealSelf | 15 |
| Healthgrades | 15 |
| Toptal | 14 |
Yelp and Expertise.com are the two directories AI engines trust most when advising buyers. BBB comes third. For B2B categories, Clutch and G2 dominate. For health, RealSelf and Healthgrades. For home services, Angi and HomeAdvisor.
The pattern: AI engines recommend vertical-specific directories over general ones. A buyer asking about cosmetic surgery gets pointed to RealSelf. A buyer asking about managed detection and response gets pointed to Gartner and G2.
Which Categories Are Most Visible to AI?
We scored each category on a 0-100 scale using four signals: entity naming rate (40% weight), source/directory citation rate (35%), cross-engine brand agreement (15%), and structured output rate (10%).
The top 10:
| Rank | Category | Vertical | Score |
|---|---|---|---|
| 1 | UI/UX Design Agencies | Professional Services | 56 |
| 2 | WordPress Development Agencies | Tech & SaaS | 54 |
| 3 | Virtual CISO Services | Professional Services | 54 |
| 4 | Shopify Development Agencies | Tech & SaaS | 53 |
| 5 | Mobile App Development | Tech & SaaS | 52 |
| 6 | Translation Services | Professional Services | 50 |
| 7 | Accounting & CPA Firms | Professional Services | 50 |
| 8 | Electricians | Local & Home Services | 50 |
| 9 | Moving Companies | Local & Home Services | 50 |
| 10 | AI SEO & AEO Agencies | Tech & SaaS | 49 |
Professional Services dominated. Seven of the top 15 categories belong to this vertical. Agencies, consultancies, and professional firms produce the kind of structured online presence (case studies, client lists, review profiles) that AI engines parse well.
The bottom 10:
| Rank | Category | Vertical | Score |
|---|---|---|---|
| 41 | Pest Control | Local & Home Services | 42 |
| 42 | Managed Detection & Response | Tech & SaaS | 40 |
| 43 | Headless Commerce Platforms | Tech & SaaS | 40 |
| 44 | Telecommunications Providers | Traditional B2B | 40 |
| 45 | Mental Health Services | Health & Wellness | 40 |
| 46 | Wholesale Trade Distributors | Traditional B2B | 39 |
| 47 | Regenerative Medicine | Health & Wellness | 39 |
| 48 | Home Health Care | Health & Wellness | 38 |
| 49 | Industrial Goods Manufacturers | Traditional B2B | 37 |
| 50 | Addiction Treatment | Health & Wellness | 36 |
Healthcare and Traditional B2B filled the bottom. Addiction treatment centres, home health care providers, and industrial manufacturers scored lowest. These industries have weaker structured data footprints online: fewer review profiles, less schema markup, and more fragmented web presence.
The Vertical Gap
| Vertical | Avg Score |
|---|---|
| Professional Services | 50 |
| Tech & SaaS | 47 |
| Local & Home Services | 47 |
| Traditional B2B | 43 |
| Health & Wellness | 43 |
Professional Services outperformed everything else. The 7-point gap between Professional Services (50) and Traditional B2B (43) represents a real structural difference in how these industries present themselves online.
Tech companies and local service providers tied at 47. HVAC companies, electricians, and dog trainers scored nearly as well as Shopify agencies and cloud platforms. Why? Because local services have Yelp profiles, BBB listings, and Angi pages that AI engines pull from. Traditional B2B companies (fleet management, wholesale distributors, industrial furniture) often lack these public-facing profiles entirely.
80+ Brands Named by All Three Engines
Cross-engine consensus is the strongest signal in this study. If Claude, ChatGPT, and Perplexity all independently name the same business, that brand has achieved genuine AI visibility.
Highlights from the consensus list:
| Brand | Category |
|---|---|
| McKinsey, BCG, Bain, Deloitte, PwC, EY, KPMG | Management Consulting |
| CBRE Group, JLL | Commercial Real Estate |
| Steelcase, Herman Miller, Haworth, Knoll, HON | Industrial Office Furniture |
| Waste Management, Republic Services, Clean Harbors | Waste Management |
| Geotab, Samsara, Fleetio, Verizon Connect | Commercial Fleet Management |
| Teladoc Health, Amwell, MDLive, Doctor on Demand | Telehealth |
| Planet Fitness, Anytime Fitness, LA Fitness | Fitness & Gyms |
| Orkin, Terminix, Rentokil, Arrow | Pest Control |
| Stanley Steemer, Chem-Dry, SERVPRO | Carpet Cleaning |
| PODS, Allied Van Lines | Moving Companies |
| WebDevStudios, Human Made | WordPress Development |
| Cover Genius | Embedded Insurance |
Traditional B2B had the strongest consensus. Every AI engine names Steelcase for office furniture, Waste Management for waste services, and Geotab for fleet management. These companies own their categories so completely that AI has no choice but to name them.
The opposite was true for fragmented markets. Dog trainers and driving schools had near-zero cross-engine agreement. No single brand dominates these categories nationally, so each AI engine picks different local or regional names.
What This Means for Business Visibility
AI engines pull brand names from structured data: review site profiles, directory listings, schema markup, and consistent NAP (Name, Address, Phone) data across the web. The businesses that scored highest in this study share three traits.
First: they appear on vertical-specific directories. A Shopify development agency on Clutch gets cited. One that only has a website does not. Second: they have review volume on platforms AI engines trust (Yelp, G2, BBB, Expertise). Third: they have structured data on their own websites (schema markup, clear service pages, location data).
The businesses at the bottom of this study are invisible not because they're small. Industrial Goods Manufacturers is a $4 trillion global industry. It scored 37 because the companies in it don't maintain the public digital signals AI engines look for.
The Gemini Question
Gemini 2.0 Flash returned empty responses across all 250 queries. Zero entities. Zero sources. Zero directory mentions.
We ran the same prompts through the same API format used for the other three engines. Gemini's API returned valid HTTP responses with empty content fields. We re-ran all 250 queries with a second script. Same result.
We're publishing this as-is. Whether this reflects a technical issue with Google's API, a content policy that blocks business recommendations, or a model limitation, the observable result is the same: Gemini produced nothing usable for this study. We'll re-test in the Q4 edition.
Methodology
750 queries total. 250 per engine (Claude, ChatGPT, Perplexity). 50 categories, 5 query templates per category. All API calls made within a single 48-hour window. Claude and ChatGPT ran at temperature 0 for deterministic output. All queries in isolated sessions with no context carryover.
Responses were auto-parsed for: brand names (regex extraction from markdown-formatted text), URLs, organic directory mentions (checked against a list of 40 known directories), response type classification (entity list, comparison table, narrative, refusal, or empty), and structured output detection.
Scoring: Entity naming rate (40%), source/directory citation rate (35%), cross-engine brand agreement (15%), structured output rate (10%).
Limitations: AI responses are non-deterministic even at temperature 0. Training data cutoffs differ across engines. Perplexity uses live search while Claude and ChatGPT rely on training knowledge. Sample size of 5 queries per engine per category is directional, not statistically rigorous. This measures what AI names today. It does not predict future citations.
Download the Data
Full dataset: 750 query responses with parsed brands, URLs, directories, scores, and metadata.
Download CSV (750 rows) · Study Summary (Markdown)
Licensed under CC BY 4.0. Cite as: "DirJournal AI Citation Study Q3."
Frequently Asked Questions
Which AI engine cites the most businesses?
Which directories do AI engines recommend most?
Which business categories are most visible to AI search engines?
Join the DirJournal newsletter
Weekly insights on directories, listings, SEO, and how businesses get found online.
No spam. Unsubscribe anytime.
Found this useful?
Share this article
Recommended for You

Free AI Visibility Scanners: 12 Tools That Check Your Brand on ChatGPT, Gemini & Perplexity
An honest comparison of 12 AI visibility scanners — what is actually free, what is paywalled, and…

Personal Injury Lawyers Spend the Most on Ads But Score Lowest in AI Search
Personal Injury scored 71. Estate Planning scored 81. The legal specialty with the highest Google…

ChatGPT vs. Perplexity vs. Gemini: Which LLMs Are Driving Real Conversions?
The largest academic study to date contradicts the agency narrative about AI-driven conversions.…
Related Resources
Looking for verified service providers? Browse our directory categories below — all human-audited and trusted by decision-makers since 2007.