Independent benchmark · 2026 Q3

Which airline brands make the AI shortlist?

The U.S. Airline AI Visibility Index measures which brands four major answer engines recommend for common travel decisions—and publishes the questions, scoring rules, uncertainty, and limitations behind every result.

The 27 July 2026 release includes 240 successful observations. Recommendation-decision agreement was 96.7% before 14 disputed observations were resolved through disclosed GPT-5 adjudication rather than human review.

  • 10 airline brands
  • 20 frozen buyer questions
  • 4 answer engines
  • 240 collected responses

Release record

U.S. Airline AI Visibility Index

A repeatable benchmark of which airline brands AI answer engines recommend for common domestic-travel decisions.

Results published · AI-adjudicated2026 Q3 release · protocol 1.0.3
Included brands
10
Buyer questions
20
Answer engines
4
Observed responses
240

Results status

Reviewed rankings and intervals.

RankBrandIndex score95% intervalRecommendation shareFirst-choice shareStabilityCitation support
1Delta Air Lines64.252.8–75.479.2%29.2%86.7%77.4%
2Southwest Airlines61.247.7–73.574.2%30.8%86.7%74.7%
3Alaska Airlines46.135.5–56.962.1%8.8%83.3%71.1%
4JetBlue Airways39.726.8–53.150.4%14.6%83.3%69.4%
5American Airlines39.226.5–51.852.9%7.1%85.8%71.7%
6United Airlines38.629.5–48.054.6%1.3%76.7%59.5%
7Frontier Airlines14.95.3–25.721.3%0.0%93.3%54.9%
8Spirit Airlines13.34.7–23.317.5%3.3%94.2%57.1%
9Hawaiian Airlines10.53.5–20.613.3%3.8%86.7%75.0%
10Allegiant Air2.60.0–6.13.8%0.0%97.5%88.9%

Headline formula

Index score = (0.70 × recommendation share) + (0.30 × first-choice share)

Average the three runs within each prompt-engine pair, then average prompt-engine pairs with equal weight. Report engine-level results alongside the combined score. Report 95% prompt-level bootstrap intervals using 10,000 resamples. Treat overlapping intervals as statistically unresolved rather than declaring a definitive winner.

Frozen comparison set

Ten consumer-facing airline brands

  1. 01Southwest Airlines
  2. 02American Airlines
  3. 03Delta Air Lines
  4. 04United Airlines
  5. 05Alaska Airlines
  6. 06JetBlue Airways
  7. 07Frontier Airlines
  8. 08Spirit Airlines
  9. 09Allegiant Air
  10. 10Hawaiian Airlines

Include the ten highest-volume consumer-facing U.S. airline brands in calendar-year 2025 T-100 passenger data after excluding regional operators that do not sell itineraries under their own public brand. Review the federal source

Equal engine weight

Four answer engines

  • ChatGPTOpenAI

    Record the provider-visible product or model label, retrieval state, response or conversation identifier when exposed, and timestamp.

  • ClaudeAnthropic

    Record the provider-visible product or model label, retrieval state, response or conversation identifier when exposed, and timestamp.

  • GeminiGoogle

    Record the provider-visible product or model label, grounding state, response or conversation identifier when exposed, and timestamp.

  • PerplexityPerplexity AI

    Record the provider-visible product or model label, search state, response or conversation identifier when exposed, and timestamp.

Frozen question set

Every question is public before collection.

Wording, order, and category membership remain unchanged during fieldwork.

Network and trip fit5 questions
  1. network-01Which U.S. airlines should I compare for a cross-country trip?
  2. network-02Which U.S. airline has a strong route network for West Coast travelers?
  3. network-03Which U.S. airline has a strong route network for East Coast travelers?
  4. network-04Which U.S. airline should I compare for flights to Hawaii?
  5. network-05Which airline should I compare for flights between California and New York?
Price and flexibility5 questions
  1. value-01Which airline should I consider for a low-cost domestic trip after fees are included?
  2. value-02Which U.S. airline is a good choice for travelers who want free checked bags?
  3. value-03Which airline is a good choice for a last-minute domestic booking?
  4. value-04Which U.S. airline is a good choice for travelers who want flexible ticket changes?
  5. value-05Which airlines offer good value for an occasional domestic traveler?
Traveler experience5 questions
  1. experience-01Which U.S. airline is best for a family traveling with checked bags?
  2. experience-02Which airline is a good choice for travelers who value schedule reliability?
  3. experience-03Which airline is a good choice for travelers who value customer service?
  4. experience-04Which airline is a good choice for travelers who prefer more legroom in economy?
  5. experience-05Which U.S. airline is a good choice for travelers with accessibility needs?
Loyalty and business travel5 questions
  1. loyalty-01Which airline is a good choice for frequent business travel within the United States?
  2. loyalty-02Which U.S. airline loyalty program is most useful for occasional travelers?
  3. loyalty-03Which airline loyalty program is most useful for frequent domestic travelers?
  4. loyalty-04Which airline offers a good domestic premium-cabin experience?
  5. loyalty-05Which airlines should a small business compare for frequent employee travel?
Machine-readable release

Download the protocol, stable prompt and carrier IDs, observation schema, scoring rules, results, controls, and limitations.

Download JSON

The evidence

Four views of the same shortlist.

The headline score stays narrow. Supporting measures show whether a brand appears consistently and whether the answer provides evidence.

Recommendation share

The percentage of eligible responses that explicitly present a brand as a suitable choice or shortlist candidate.

First-choice share

The percentage of responses where a brand is the first affirmative recommendation for the buyer’s question.

Stability

How often the recommendation repeats across the three identical runs, reported by prompt and answer engine.

Citation support

How often an attached source directly supports the brand claim. This is reported separately and does not influence rank.

Coding agreement

Forty-eight observations were double-coded. Pre-adjudication agreement was 96.7% for brand recommendation decisions, 100% for first choice, and 70.8% for complete citation-support sets.

Why this market

A category where the shortlist has consequences.

Air travel combines recognizable brands, frequent comparison questions, public market data, changing fees, and decisions where a generic answer can hide important tradeoffs.

Objective inclusion

Brands are selected before answers are collected.

The comparison set comes from federal passenger data, with a published rule for excluding regional operators that do not sell trips under their own public brand.

Review the BTS source

Buyer relevance

Questions reflect real travel decisions.

The frozen set covers route fit, total trip value, flexibility, family travel, reliability, accessibility, loyalty, and business travel.

Release process

Publish the rules first. Collect the answers second.

A result is useful only when readers can see what was fixed in advance and what changed later.

Freeze

Version the brands, prompts, engines, formula, classification rules, fieldwork window, and exclusions before the first answer is collected.

Collect

Run every prompt three times per engine in fresh sessions and preserve raw answers plus provider metadata.

Code and check

Classify recommendations with written rules, double-code a sample, and record disagreements.

Publish

Release combined and engine-level results, uncertainty intervals, source patterns, raw-data notes, and limitations.

For journalists and analysts

Built to be inspected, quoted, and challenged.

Every edition will retain a stable release record so reporting can distinguish measured findings from interpretation.

01

Public protocol

Exact questions, brand-selection rule, scoring formula, collection requirements, and exclusions remain available with the release.

02

Reproducible tables

Results will show denominators, engine splits, confidence intervals, missing observations, and release-version notes.

03

Evidence boundaries

The index reports recommendation frequency. It does not claim that a recommendation is correct, safe, or suitable for every traveler.

Interpretation

Read movement as evidence, not destiny.

Models, retrieval systems, public information, and travel conditions change. Each edition is a dated sample. Comparisons across editions will use the same core question set and disclose every protocol change.

What the index can show

Which included brands appear, how prominently, on which engines, for which questions, and how consistently during the collection window.

What it cannot prove

Market-wide consumer behavior, service quality, safety, financial performance, or the right airline for an individual itinerary.

Review limitation

The 14 disputed observations were adjudicated by GPT-5 at the owner’s request, not by a human reviewer. The dated change log retains that departure from the planned method.

Research enquiries

Request the protocol, fieldwork notes, or an interview.

Journalists, analysts, airlines, and researchers can ask about the frozen method, reviewed findings, or release limitations. Questions do not influence inclusion, scoring, or corrections.

Index FAQ

Questions about the benchmark

Are the airline rankings available yet?

Yes. The 2026 Q3 rankings are published with 95% prompt-level bootstrap intervals. Because those intervals overlap for several brands, close rank positions should be treated as statistically unresolved.

Can an airline pay to be included or improve its score?

No. Inclusion follows the published federal-data rule, and the prompts, engine weights, and formula are frozen before collection. Sponsorship does not affect the benchmark.

Does the top-ranked airline provide the best service?

No. The index measures how often selected AI answer engines recommend an included brand for the frozen questions. It does not measure safety, operational quality, fares, or individual suitability.

Why run each question three times?

Repeated runs show whether a recommendation is stable or appeared once. The three runs are averaged before each prompt-engine pair contributes to the combined result.

What model details will be disclosed?

Every observation will report the provider-visible product or model label, retrieval or grounding state, timestamp, and response or conversation identifier when the product exposes one. Unavailable metadata will be marked null rather than inferred.

How will later editions remain comparable?

The core prompt set and scoring formula remain stable. Any change to brands, models, prompts, classification, or fieldwork is versioned and disclosed before the next collection begins.