Skip to content

Blog

What we discovered when testing AI hotel recommendations

By Blair Roche

What we discovered when testing AI hotel recommendations

Search marketers have spent the last two decades optimizing for defined, repeatable search behavior. Traditional hotel search is easy to recognize and, therefore, relatively easy to strategize around: a traveler inputs narrow parameters into search boxes or dropdowns, like location, brand, price, and number of guests.

AI gives travelers a different way to express their needs, but instead of translating a trip into keywords and filters, travelers simply explain their trip to the LLM conversationally and likely add much more context. Instead of inputting keywords like ‘Chicago hotels,’ ‘McCormick Place,’ ‘Marriott Chicago,’ or ‘hotels under $250,’ consider a prompt from our recent research:

“I’m traveling to Chicago for work. I’ll be there for two nights and want to stay somewhere convenient for meetings, restaurants, and getting around.”

Now, hotels must expand optimization from a traditional search query to a description of a decision. The traveler has given the AI information about the purpose of the trip, length of stay, priorities, location needs and implicitly, the type of hotel experience that would work for them. The traveler can then add another layer by telling the AI that the traveler has a $200 budget or is a Marriott Bonvoy member.

We spent a week trying to understand this shift so that hotels can better optimize their search efforts for visibility and consideration in AI-generated results, with our initial test using ChatGPT. This is a first look at one model’s behavior. We picked ChatGPT because it’s the most widely used consumer assistant with live web search, and we’ll expand to other platforms in later phases.

Here’s what we found.

The test: Pilot study for AI hotel discovery

We began our test by building a controlled hotel recommendation benchmark using ChatGPT with web search. 

The study included 400 different travel scenarios, with variables including destination, traveler type, budget and loyalty preference. Each scenario was run five times so we could also measure recommendation consistency.

In total, the benchmark produced 2,000 completed responses and 10,000 individual hotel recommendations.

For each response, we captured the exact prompt, hotels recommended, recommendation order, explanation provided, supporting domains and underlying response data. Instead of trying to infer how ChatGPT’s recommendation system works, we measured its outputs and based our analysis on those observable results.

The prompts were intentionally conversational. But the results that emerged were less like a traditional search ranking and more like a series of context-dependent consideration sets.

Observation #1: There may not be one AI ranking for a hotel

One of the clearest signals was how much the recommendation set moved. When we ran the exact same prompt five times, average overlap among the recommended hotels was only 37.3%

Even the hotel most frequently ranked first for a given scenario held that position in an average of only 62.4% of runs. Across the entire benchmark, only one prompt produced the exact same five-hotel recommendation set in all five runs. 

Some of that variation may reflect the nature of live web search itself. Different runs can retrieve or prioritize different sources, meaning the 37.3% overlap should not be interpreted as a measure of model instability alone. Our test was designed to measure the recommendation experience a traveler actually receives, where model reasoning and live information retrieval work together.

Traditional search has trained marketers to think about position, with marketers focusing their efforts on ranking for a particular query. In our test, the recommendation experience was highly fluid, even when the traveler’s prompt remained exactly the same. The important question for marketers may be less about whether a hotel is “ranked #1” and more about how frequently it enters the consideration set for a particular type of traveler.

Observation #2: Visibility shifts with the traveler

There wasn’t one set of brands or properties that consistently dominated the results. For these comparisons, we held the other prompt variables constant and changed one element at a time. As we changed the context of the trip, including who was traveling, what they wanted to spend, and whether they had a loyalty preference, the results shifted as well (within our benchmark):

  • Changing traveler type changed roughly 73% of the recommendation set
  • Changing budget changed about 85%
  • Adding a loyalty preference changed roughly 96%

We also saw a difference between being frequently considered and being the top choice. Some properties appeared across a wide range of scenarios, while others surfaced less often but were more likely to earn the #1 recommendation when the circumstances were right.

For marketers, that means AI visibility is probably not one number. Presence, frequency, position, and the traveler context that triggered the recommendation all matter.

The bigger opportunity is understanding which travel contexts your property is relevant for and where you are missing from consideration altogether. That means hotels may need to think about AI visibility at the scenario level. A downtown property, for example, might evaluate whether it appears for business travelers, families, loyalty members, budget-conscious travelers, event attendees, and other relevant trip contexts—not simply whether it appears for a broad query like “hotels in Chicago.”

Observation #3: The information ecosystem matters 

The sources cited alongside those recommendations give us another view into how AI discovery differs from traditional search.

Across the benchmark, roughly 44% of the citations we observed pointed to direct hotel or brand websites, compared with approximately 17% pointing to OTAs. Metasearch, editorial publications, review platforms and other sources made up the rest.

That does not tell us that one source type “drives” an AI recommendation. Our benchmark measures observable behavior, not the model’s internal ranking logic.

It does tell us that first-party hotel content remains very much part of the information ecosystem AI uses when helping travelers make decisions.

An assistant can synthesize information from the hotel, an OTA, a publisher, a review platform and other sources before presenting a much smaller set of options back to the traveler. That makes AI visibility partly an information-availability problem.

Hotels cannot control every source an AI assistant encounters, but they can influence whether clear, current and differentiated information about their properties exists across the ecosystem. The question is both “Can a search engine find my property?” and “Can an AI assistant understand when my property is the right answer?” That puts renewed importance on information such as amenities, location context, loyalty benefits, traveler fit, property differentiators and other details that help explain why a hotel fits a particular trip.

Next steps: From keywords to decisions

Our most important takeaway isn’t any single visibility metric, but how different the underlying consumer behavior is.

A traveler no longer has to know how to turn their needs into search terms, select filters, open ten tabs and assemble the answer themselves. They can describe the trip and let the AI do some of that work.

Marketers need to shift their strategy from search ranking to understanding: “Under what traveler circumstances does our property become part of the consideration set?”

That creates a different set of strategic questions around content, distribution, loyalty, reputation, publisher presence, pricing, and the information AI can access about a property.

We are still early in understanding how those pieces fit together, and this first phase measures observable recommendation behavior rather than claiming to explain the underlying algorithm. But even these first findings point to a meaningful change in the discovery experience. Hotels should begin measuring not only how they rank in search, but when and how often they enter AI-generated consideration sets, for which travelers, and in which contexts.

Why this matters now

This is an early benchmark, and there is much more to test across models, destinations and traveler scenarios. But hotels do not need to wait for a definitive study of AI ranking systems to start asking a different question.

Traditional search gave marketers a relatively visible playing field: queries, rankings, impression share and clicks. AI discovery may create a much less predictable consideration set, assembled around the specific needs of an individual traveler.

That makes understanding whether, when and why a property enters that consideration set a new visibility problem for hotel marketers. While the mechanics will evolve, marketers should start planning for the strategic shift occurring today.

What Koddi will measure next 

This initial test is the first phase of a broader Koddi research effort into how AI is changing travel discovery and commerce. It was focused on the recommendation layer: what hotels enter the consideration set, how stable those recommendations are, what sources support them and how traveler context changes the outcome.

Next, we are looking at how AI platforms guide a traveler after the initial question. That includes how they encourage travelers to narrow their choices, which next actions they suggest, when hotel and booking options begin to appear, what sources they rely on and where commercial experiences enter the conversation.

The larger question is how AI changes the full path from inspiration to decision to transaction. Phase 1 measured only the entry point. Understanding what happens after that initial recommendation—and how AI reshapes the path toward booking—may ultimately be even more important for the travel industry. 

Ready to get started?