Understanding the Move from Accuracy V1 to Accuracy V2
For Hotel & Resort Brands A plain-language guide as to why your AI Accuracy score is changing — and what it does (and doesn't) mean.
|
In one sentence Bonafide has upgraded the approach taken to grade AI answers — from a strict right/wrong test scale to a more forgiving, partial-credit scale. For hotels specifically, this has meant a modest, fairly predictable dip in reported Accuracy — not because AI platforms know less about your property, but because the new grading is stricter about single wrong facts, which is what most hotel questions come down to. |
Why you're seeing this
Since Bonafide launched, your Accuracy score has been calculated using what we now call Accuracy V1. Starting with June 2026 data, we moved to a new methodology, Accuracy V2. Beginning in July 2026, V1 is being retired for all new data — your historical V1 scores stay exactly as they were, but new scores going forward will be calculated using the V2 model.
In the platform, you can toggle between V1 and V2 for your recent data so you can see both side by side. This article explains, in plain terms, what actually changed, why your number may have moved up or down even though nothing changed about your listings or content, and what we recommend if you want to compare V1 and V2 numbers fairly.
The one-sentence version of each methodology
Think of it like the difference between a true/false quiz and an essay graded for partial credit.
Accuracy V1 — the true/false quiz
Every time an AI platform (ChatGPT, Gemini, Claude, Perplexity, Meta Llama, etc.) answered a question about your hotel, that answer was marked simply right or wrong — a green check or a red X. There was no partial credit. An answer that was 90% correct scored exactly the same (zero) as an answer that was completely made up.
Accuracy V2 — the partial-credit essay
Every answer is now read by an AI “grader” that compares it to your verified facts and assigns one of five grades, from Bad (0 stars) to Excellent (4 stars), based on how much of the answer was correct and whether it contradicted the facts outright. An answer that gets most of the details right but misses one small item is no longer treated the same as an answer that is completely wrong. Partial credit is now given using a ‘star’ scale system.
The V2 star scale
|
Stars |
Label |
What it means |
|---|---|---|
|
4 |
Excellent |
All key facts captured; nothing contradicted |
|
3 |
Good |
Most key facts right; minor supporting detail may be missing |
|
2 |
Neutral |
Some key facts right, some missing — partially correct |
|
1 |
Poor |
Few key facts right; vague or incomplete |
|
0 |
Bad |
No key facts right, or the answer contradicts your truth |
A response needs at least 2 stars (Neutral or better) to be considered directionally correct — meaning it captured real facts and didn't contradict your source of truth. Anything scoring below that (Poor or Bad) is treated as a miss, the same value that Accuracy V1 would have considered “incorrect.” We come back to why this 2-star line matters later in this article.
Side-by-side: how the math actually works
|
Accuracy V1 |
Accuracy V2 |
|
|---|---|---|
|
Grades each answer as |
Correct or incorrect (binary) |
0–4 stars (five levels) |
|
Who grades it |
Structured correct/incorrect check |
AI grader using a guided rubric |
|
Handles an “almost right” answer |
No partial credit — scored as wrong |
Yes — scored as Neutral or Good |
|
Handles paraphrasing / unit conversions |
Poorly |
Yes |
|
Score formula |
Correct answers ÷ Total answers |
Sum of stars earned ÷ Total stars possible |
|
Content gaps (“nulls”) |
Excluded from the score |
Excluded from the score |
Worked example: Five AI platforms answer the same question. Under V2 they score 2, 2, 2, 4, and 4 stars. The math is 14 stars earned out of 20 possible (5 platforms × 4 stars) — a 70% score. Under V1, only the two answers that scored a full 4 would likely have counted as “correct,” and the rest — even though they captured real, partially-correct information — would have scored as a flat “incorrect” score. That difference in how partial answers are treated is the reason why Accuracy Scores will move.
Why did MY score move — and why did it move down?
This is the part that matters most for your business, so it's worth an analogy. Imagine your amenities page is being read aloud by different guests, and someone is scoring how well each guest reports it back:
- Under the old scoring (V1), a guest who got the room count, the pool hours, and the pet policy right but mixed up one number — say, they said the gym opens at 6am instead of 5am — receives the same zero (Incorrect value) as a guest who invented a spa that doesn't exist at the property.
- Under the new scoring (V2), that first guest still loses points for a mistake, but is recognized for getting most of it right. The second guest, who invented facts, is still penalized to the bottom of the scale because V2 is strict about outright contradictions.
Hotel and resort questions are usually built around one exact, specific fact: an exact room count, a specific check-in cutoff, whether a named accessibility feature exists, whether a resort fee applies. There's rarely a “partially right” version of “does this hotel have a roll-in shower” — an AI either knows the specific fact or it doesn't. When it gets one of these specific facts wrong, V2's stricter no-contradiction rule counts that as a real miss the same way V1 did. That's why hotel scores, across the board, have trended modestly lower under V2 — it isn't a sign your content got worse or that AI platforms suddenly forgot about you.
One thing worth knowing: destination and DMO content (which asks more list-style questions, like naming local festivals or markets) has actually trended the positive upward direction under V2, because partial credit rewards a mostly-complete list in a way it never rewarded a mostly-complete hotel fact. If you ever see a company-wide Bonafide benchmark that blends every vertical together, don't use it to judge your own hotel's movement — always compare your score to other hotels, not to a mixed bag that includes destinations.
What the data actually shows (June 2026, Hotels only)
We compared V1 and V2 scores calculated from the same June 2026 data, across every active hotel and resort customer.
|
Metric |
Value |
|---|---|
|
Average V1 Accuracy |
54.4% |
|
Average V2 Accuracy |
47.9% |
|
Average change |
–6.6 points |
|
Range across hotel brands measured |
–2.1 to –9.9 points — every brand moved the same direction |
|
Hotel brands sampled |
11 |
|
Key takeaway Every hotel brand in our sample moved in the same direction (down) under V2, within a fairly tight band (–2.1 to –9.9 points). That consistency is actually reassuring: it means this is a predictable, expected effect of the new methodology — not something specific to your property. |
A universal effect, on top of the vertical effect
We also compared the shift by AI platforms (not by customer). Every major model — Anthropic Claude, Google Gemini, Meta Llama, OpenAI's ChatGPT, and Perplexity — moved downward under V2 by somewhere between 2 and 6.5 points on average, regardless of vertical. That tells us part of the shift is a consistent, expected effect of V2's stricter contradiction penalty across every model. The additional swing on top of that — up for some verticals, down for others — is being driven by the type of content (specific single facts vs. lists) being tested.
Can we normalize or “handicap” the comparison?
Yes — and for hotels specifically, the data supports doing this with reasonable confidence, because the shift is consistent in direction and size across brands.
Recommended: a hotel-specific offset
|
Vertical |
Suggested offset |
How to use it |
Confidence |
|---|---|---|---|
|
Hotels |
+6.6 points |
Add to a V2 score to estimate its V1-equivalent |
Moderate-to-good — tight, consistent spread (±2–3 pts), all 11 sampled brands moved the same direction |
Estimated V1-equivalent (Hotels) = V2 score + 6.6
We'd suggest this offset is reasonably safe to use in portfolio-level trend charts and executive summaries, with a plus-or-minus few points of margin noted. It's based on one month of paired data across 11 hotel brands, so we recommend recalculating it each quarter as more paired V1/V2 data accumulates — the estimate will get more precise, not less, over time. It should be noted that Accuracy V1 will only be available through July 2026 (June 2026 data), so there are only a couple months of overlap between the two measurement types.
A more exact complementary option: the “2-Star Comparable Rate”
If you want a bridge number that isn't an estimate at all, there's a more exact option available directly from the V2 grading data. Since a response needs at least 2 stars (Neutral or better) to be considered directionally correct, we can recompute a true binary “percent correct” figure from V2's underlying grades — the same yes/no logic V1 used, just applied to V2's more careful grading:
2-Star Comparable Rate = (responses graded 2 stars or higher) ÷ (total scored responses)
This is a real recount of your own data, not a group estimate. We recommend it if you want an exact apples-to-apples read against your own V1 history rather than a portfolio-wide adjustment.
|
Bottom line for your team 1) Your V2 score measures something genuinely more nuanced than V1 did — a lower V2 number does not mean AI suddenly knows less about your property. 2) Compare your score to other hotels, not to a blended average that includes other verticals — destinations tend to move the opposite direction. 3) For a single comparable trend line, add 6.6 points to your V2 score as a reasonable estimate of its V1-equivalent, or use the 2-Star Comparable Rate for an exact bridge. 4) Your historical V1 data isn't going anywhere — it stays exactly as recorded. Only new data going forward uses V2. |
Frequently asked questions
My score dropped after the switch. Did AI platforms get worse at describing my property?
Almost certainly not. A drop is far more likely to reflect V2's stricter treatment of any single wrong or contradicted fact — something hotel questions are especially likely to hinge on, since they usually have one exact correct answer (a room count, a cutoff time, whether a feature exists).
Will you go back and recalculate my old scores under V2?
No. Recomputing historical answers under V2 would require re-running the original AI conversations, which can't be reliably reconstructed after the fact. Your historical V1 data stays exactly as it was recorded.
Which number should I actually track going forward?
Track Accuracy V2 as your primary, ongoing metric — it's a more informative measure of what AI is actually telling travelers, including where an answer was close but not perfect. Use the +6.6 offset or the 2-Star Comparable Rate only when you specifically need to compare against pre-July history.
Is a -6.6 point drop something I should escalate or worry about?
Only if your own drop is meaningfully larger than that (for example, 15–20+ points) or outside the –2 to –10 point band we're seeing across hotels — that would be worth flagging to your Bonafide contact so we can check whether something specific to your content changed, versus the expected methodology effect everyone else is seeing.
How long will Accuracy V1 remain visible in the Bonafide Reporting UI?
Accuracy V1 will be presented through July 2026, after which only Accuracy V2 will be shown in reporting.