Recommendation share
The percentage of eligible responses that explicitly present a brand as a suitable choice or shortlist candidate.
Independent benchmark · 2026 Q3
The U.S. Airline AI Visibility Index measures which brands four major answer engines recommend for common travel decisions—and publishes the questions, scoring rules, uncertainty, and limitations behind every result.
The 27 July 2026 release includes 240 successful observations. Recommendation-decision agreement was 96.7% before 14 disputed observations were resolved through disclosed GPT-5 adjudication rather than human review.
Release record
A repeatable benchmark of which airline brands AI answer engines recommend for common domestic-travel decisions.
Results status
| Rank | Brand | Index score | 95% interval | Recommendation share | First-choice share | Stability | Citation support |
|---|---|---|---|---|---|---|---|
| 1 | 64.2 | 52.8–75.4 | 79.2% | 29.2% | 86.7% | 77.4% | |
| 2 | 61.2 | 47.7–73.5 | 74.2% | 30.8% | 86.7% | 74.7% | |
| 3 | 46.1 | 35.5–56.9 | 62.1% | 8.8% | 83.3% | 71.1% | |
| 4 | 39.7 | 26.8–53.1 | 50.4% | 14.6% | 83.3% | 69.4% | |
| 5 | 39.2 | 26.5–51.8 | 52.9% | 7.1% | 85.8% | 71.7% | |
| 6 | 38.6 | 29.5–48.0 | 54.6% | 1.3% | 76.7% | 59.5% | |
| 7 | 14.9 | 5.3–25.7 | 21.3% | 0.0% | 93.3% | 54.9% | |
| 8 | 13.3 | 4.7–23.3 | 17.5% | 3.3% | 94.2% | 57.1% | |
| 9 | 10.5 | 3.5–20.6 | 13.3% | 3.8% | 86.7% | 75.0% | |
| 10 | 2.6 | 0.0–6.1 | 3.8% | 0.0% | 97.5% | 88.9% |
Headline formula
Index score = (0.70 × recommendation share) + (0.30 × first-choice share)Average the three runs within each prompt-engine pair, then average prompt-engine pairs with equal weight. Report engine-level results alongside the combined score. Report 95% prompt-level bootstrap intervals using 10,000 resamples. Treat overlapping intervals as statistically unresolved rather than declaring a definitive winner.
Frozen comparison set
Include the ten highest-volume consumer-facing U.S. airline brands in calendar-year 2025 T-100 passenger data after excluding regional operators that do not sell itineraries under their own public brand. Review the federal source
Equal engine weight
Record the provider-visible product or model label, retrieval state, response or conversation identifier when exposed, and timestamp.
Record the provider-visible product or model label, retrieval state, response or conversation identifier when exposed, and timestamp.
Record the provider-visible product or model label, grounding state, response or conversation identifier when exposed, and timestamp.
Record the provider-visible product or model label, search state, response or conversation identifier when exposed, and timestamp.
Frozen question set
Wording, order, and category membership remain unchanged during fieldwork.
Download the protocol, stable prompt and carrier IDs, observation schema, scoring rules, results, controls, and limitations.
The evidence
The headline score stays narrow. Supporting measures show whether a brand appears consistently and whether the answer provides evidence.
The percentage of eligible responses that explicitly present a brand as a suitable choice or shortlist candidate.
The percentage of responses where a brand is the first affirmative recommendation for the buyer’s question.
How often the recommendation repeats across the three identical runs, reported by prompt and answer engine.
How often an attached source directly supports the brand claim. This is reported separately and does not influence rank.
Forty-eight observations were double-coded. Pre-adjudication agreement was 96.7% for brand recommendation decisions, 100% for first choice, and 70.8% for complete citation-support sets.
Why this market
Air travel combines recognizable brands, frequent comparison questions, public market data, changing fees, and decisions where a generic answer can hide important tradeoffs.
Objective inclusion
The comparison set comes from federal passenger data, with a published rule for excluding regional operators that do not sell trips under their own public brand.
Review the BTS sourceBuyer relevance
The frozen set covers route fit, total trip value, flexibility, family travel, reliability, accessibility, loyalty, and business travel.
Release process
A result is useful only when readers can see what was fixed in advance and what changed later.
Version the brands, prompts, engines, formula, classification rules, fieldwork window, and exclusions before the first answer is collected.
Run every prompt three times per engine in fresh sessions and preserve raw answers plus provider metadata.
Classify recommendations with written rules, double-code a sample, and record disagreements.
Release combined and engine-level results, uncertainty intervals, source patterns, raw-data notes, and limitations.
For journalists and analysts
Every edition will retain a stable release record so reporting can distinguish measured findings from interpretation.
01
Exact questions, brand-selection rule, scoring formula, collection requirements, and exclusions remain available with the release.
02
Results will show denominators, engine splits, confidence intervals, missing observations, and release-version notes.
03
The index reports recommendation frequency. It does not claim that a recommendation is correct, safe, or suitable for every traveler.
Interpretation
Models, retrieval systems, public information, and travel conditions change. Each edition is a dated sample. Comparisons across editions will use the same core question set and disclose every protocol change.
Which included brands appear, how prominently, on which engines, for which questions, and how consistently during the collection window.
Market-wide consumer behavior, service quality, safety, financial performance, or the right airline for an individual itinerary.
The 14 disputed observations were adjudicated by GPT-5 at the owner’s request, not by a human reviewer. The dated change log retains that departure from the planned method.
Research enquiries
Journalists, analysts, airlines, and researchers can ask about the frozen method, reviewed findings, or release limitations. Questions do not influence inclusion, scoring, or corrections.
Index FAQ
Yes. The 2026 Q3 rankings are published with 95% prompt-level bootstrap intervals. Because those intervals overlap for several brands, close rank positions should be treated as statistically unresolved.
No. Inclusion follows the published federal-data rule, and the prompts, engine weights, and formula are frozen before collection. Sponsorship does not affect the benchmark.
No. The index measures how often selected AI answer engines recommend an included brand for the frozen questions. It does not measure safety, operational quality, fares, or individual suitability.
Repeated runs show whether a recommendation is stable or appeared once. The three runs are averaged before each prompt-engine pair contributes to the combined result.
Every observation will report the provider-visible product or model label, retrieval or grounding state, timestamp, and response or conversation identifier when the product exposes one. Unavailable metadata will be marked null rather than inferred.
The core prompt set and scoring formula remain stable. Any change to brands, models, prompts, classification, or fieldwork is versioned and disclosed before the next collection begins.