The r/MedSpa post "Before buying an AI receptionist, ask it these 7 questions" proposes something simple and, for this category of purchase, unusual: test the system with scenarios you choose rather than accept the demonstration the vendor chooses. The author lists seven calls to ask a vendor to run, using fictional details so nothing real is exposed.
A caller who does not know which appointment type to book. A caller who has seen a price that differs from the website. A caller who asks for a person, including outside business hours. A caller who needs to move an appointment, where the test is whether the change appears in the actual booking system. A caller who was promised a callback and never got one. A caller with a medical question. And a caller asking what happens to the recording and transcript of the call — who can access it, and how it is deleted.
The instruction that turns this from a nice list into a working evaluation is at the end: for each test, mark completed correctly, handed off clearly, or failed, and record how much staff cleanup was needed. A clear handoff counts as a success when staff input is genuinely required.
Why the demo is the wrong test
Vendors demonstrate systems they know how to drive, on data that has been prepared, in the presence of an engineer who can intervene. A demo answers the question "can this system do this" and the buying decision is about a different question: "what does this system do when it is wrong, alone, at 7:40 on a Friday evening."
The seven-scenario format fixes that asymmetry because it replaces the vendor's script with the buyer's. Two rules make it work.
Test against your own system, not theirs. The appointment-move scenario is only meaningful if the system is writing to the practice's real scheduling software, or to a sandbox of it. A booking created in a vendor interface and never reflected in the calendar is the most common way a successful demo becomes a failed implementation, and it is invisible unless someone checks the calendar after each test.
Score the failures, not just the successes. A system that books correctly in five scenarios and silently drops a callback in two is not a system that is seventy percent ready. It is a system that will produce a category of loss the practice will not be able to see, which is worse than a system that simply fails loudly.
The metric nobody prices: staff cleanup
The post's emphasis on counting staff cleanup is the most commercially useful part of it, because cleanup is where the promised savings quietly evaporate.
The value proposition for front-desk automation is minutes. If a practice loses calls when nobody can pick up, an automated system converts some of those calls into appointments, and the arithmetic looks obvious. What the arithmetic usually omits is what happens after a call that the system handled imperfectly. A wrong appointment type booked into a treatment room that cannot deliver it, a caller given a price that no longer applies, a message that reached a shared inbox instead of a task with an owner — each of these produces staff time to detect, diagnose and repair, and that time is spent by the most expensive people in the building during the hours when they were supposed to be doing something else.
The number to watch during a pilot is therefore not calls answered. It is minutes of staff cleanup per call, recorded by the staff, plus the count of errors discovered by staff rather than by the system. A tool that adds two minutes of cleanup per call across two hundred calls a month has consumed six hours of front-desk time to save some fraction of the calls it was hired to capture. That is not automatically a bad trade, but it is a trade, and it should be measured rather than assumed.
Four failure modes worth naming before the pilot
The silent failure. The most expensive outcome is not a wrong booking; it is a call that produced no record at all. Every call, successful or not, should leave a trace with a name attached to whatever happens next. If a caller who asked for a person is not visible as an open task with an owner, the system has created a lost lead and labelled it as handled.
The invented answer. Prices, promotions, and what a treatment includes are where a language model is most likely to produce something plausible and wrong. The post's second scenario targets exactly this, and the standard should be strict: the system either states an approved price from a source the practice controls, or it does not state a price. A caller who is quoted one number and charged another is a refund conversation and a review, and the cost is not the difference in dollars.
The clinical question. A caller asking whether a treatment is suitable given a medication or a condition should be routed, not answered, and the routing should be to someone qualified to answer. That is a scope-of-practice line as much as a service-quality one, and the test is whether the system recognizes the boundary rather than whether its answer sounds reasonable.
The record. The seventh scenario — where recordings and transcripts are stored, who can access them, and how deletion works — is the one most practices ask last and should ask first. Call recordings in a medical setting carry the same handling obligations as other patient information, and the specifics depend on the practice's jurisdiction, the vendor's data location, and the agreements in place. Ask the vendor to show the retention settings and the access list rather than to describe them, and involve whoever advises the practice on privacy before a pilot that records real patient calls.
How to run the pilot so it produces a decision
A pilot that runs everywhere at once cannot be evaluated, because when something breaks there is no baseline to compare against. Three constraints make the result readable.
Start with after-hours only. Route calls outside business hours to the system and leave daytime calls with the team. This isolates the variable, it targets the calls that are currently going entirely unanswered, and it limits the blast radius of an early failure to a window where the alternative was a voicemail nobody heard.
Define the win in advance. Pick two or three numbers before the pilot starts: booked appointments from after-hours calls, handoffs that reached a human within one business day, and cleanup minutes per call. Decide now what performance would justify expanding the system, and what would end it. A pilot without a pre-committed threshold becomes a subscription by default.
Sample the calls yourself. Listen to a random set of recordings, including the ones the system handled without escalation. Reading a transcript of a conversation the practice never knew it was having is the fastest way to find the failure mode that the vendor's dashboard does not display, because dashboards count events and do not notice absences.
The decision underneath the decision
Front-desk automation is worth buying when a practice is genuinely losing calls it could otherwise convert, and when the front desk's scarce time is better spent on the callers who need judgment. It is not worth buying to reduce headcount in a practice where the front desk is the relationship, and it is not worth buying before the practice can say how many calls it actually misses.
That last point is worth an hour before any vendor conversation. A week of logging missed calls — when they arrive, what they were about, whether the caller called back — produces the number that decides whether this purchase has a return at all. It also produces something more useful than a budget: the list of specific scenarios that matter most to this practice, which is exactly the list of tests the vendor should be asked to run.