We asked ChatGPT the same question four times. The answers never matched.
The question was simple: “Best physiotherapist in Sydney – who do you recommend?” We asked it four times in one sitting, and ChatGPT answered confidently every time.
Of the 23 businesses it named, 13 appeared once and never again. Only one was named in all four answers, and ChatGPT wrote its name four different ways doing it.
The recommendation survived. The name barely did.
Ask an AI the same question four times and the answers shift. The band is that range, and a change only counts as movement once it clears it.The method in full →
Not even consistency is consistent.
ChatGPT went 0.23 to 0.56 with nothing changed but the date. Gemini agreed 0.58 on “best in Sydney” and 0.30 on the same question three days later. Consistency moves with the question, the market and the day.
ChatGPT agreed just .23 – Sydney, 7 July.
0 = never the same answer · 1 = identical every time
The grey is the noise; the lume is the one reading the data will stand behind.
What we did
On 7 July 2026 we asked two buyer questions: the Sydney physio question above, and a near-Bondi variant. Four assistants, four times each, one sitting. 32 answers, asked the way a customer would ask.
We count businesses, not strings. The methodology covers what counts as one, and the dataset carries every name exactly as each assistant wrote it.
Ask the same question four times and the shortlist barely holds. The grid below is what each assistant returned, ask by ask.
| assistant | names per answer | businesses named | named once | agreement | named in all four |
|---|---|---|---|---|---|
| ChatGPT | 11 · 15 · 4 · 9 | 23 | 13·57% | 0.23 | SportsFit Health and Rehab (Five Dock) |
| Perplexity | 6 · 6 · 5 · 7 | 13 | 6·46% | 0.31 | Andrew Noyes – Exact Physiotherapy |
| Gemini | 9 · 10 · 8 · 7 | 13 | 4·31% | 0.58 | Align Health Collective · Evoker Premium Physiotherapy Services · PhysiCo City · Sydney Physio Clinic |
| Claude | 6 · 5 · 6 · 5 | 6 | 0·0% | 0.89 | Evoker Premium Physiotherapy · Fixio · PhysiCo City Physiotherapy (Sydney CBD) · Sydney Physio Clinic (Macquarie Street) · Sydney Physio Solutions |
The same pattern held on the second question, a suburb-level ask where the lists were shorter but no steadier.
| assistant | names per answer | businesses named | named once | agreement | named in all four |
|---|---|---|---|---|---|
| ChatGPT | 8 · 4 · 6 · 7 | 12 | 6·50% | 0.42 | ISO Physiotherapy · JW Physical Health |
| Perplexity | 3 · 3 · 3 · 3 | 5 | 2·40% | 0.53 | JW Physical Health |
| Gemini | 4 · 7 · 4 · 4 | 8 | 3·38% | 0.51 | Bondi Platinum Physio · Physio Fitness (Bondi Junction) |
| Claude | 4 · 3 · 1 · 3 | 5 | 2·40% | 0.45 | The Running Room Physiotherapy |
Named once, never again
ChatGPT answered the same question with 11 names, then 15, then 4, then 9. You can be recommendation #3 in one answer and absent from the next.
Who survived
One business came through all four of ChatGPT’s answers: SportsFit Health and Rehab, written four different ways. The variants ran from the full name with its suburb down to a bare “SportsFit”, and all four are in the dataset.
Across assistants, Sydney Physio Clinic was named in every answer by two of them, and again by Claude in the Sydney repeat three days later. Evoker came through all four asks for three separate assistant-and-day pairings.
This is not a ranking. Appearing or vanishing is the assistants’ behaviour, not a judgement of any practice.
One limit, plainly. Identity is keyed on a business website domain where the assistant links one, and on the name everywhere else. Sydney’s CBD has near-identical physio names, so name-matched counts carry that caveat.
The assistants read largely the same pages and still named different businesses. Being cited and being named are two different numbers.
| assistant | agreement on businesses named | agreement on sources cited |
|---|---|---|
| ChatGPT | 0.23 | 0.36 |
| Perplexity | 0.31 | 0.61 |
| Claude | 0.89 | 1.00 |
| Gemini | 0.58 | n/a – unresolvable redirects |
The sources hold still. The names don’t.
If one question returns 23 different recommendations in four asks, a single screenshot of an AI answer is not evidence of anything. An AI ranking published without a range is noise dressed as signal.
This is why the methodology attaches a band and a date to every reading, and why The Southlume Index publishes rates with bands rather than bare ranks.
The recommendation survived. The name barely did.
Open questions
This is a pilot, and we say so. Two questions, one city, four asks per assistant, repeated in Melbourne and again in Sydney three days later.
Melbourne ran three days after Sydney, so that comparison changed the day as well as the city. We re-ran Sydney on the same day as Melbourne to separate the two.
The answer refused to be tidy, and we publish it anyway. An assistant’s consistency isn’t stable with itself from one day to the next. A snapshot is a snapshot, even if you take four of them.
The Southlume Index is measuring these categories now.
Prompt checks are a sample; referral analytics are the ground truth of actual AI-driven visibility. A zero here is ‘not found on these prompts’, never ‘invisible’.