What actually happens across ten takes
Takes one to three are performance — you can hear it, and so can everyone else. Four to six get faster and flatter as you start rushing to finish. Around seven something changes: you stop thinking about the camera and start thinking about the explanation, and your normal speaking rhythm comes back. Takes eight to ten are the usable ones.
Almost nobody gets there in one sitting on day one, and almost everybody gets there by the third session. The doctors who conclude they are bad on camera are the ones who recorded twice and stopped.
Selling products or courses alongside the clinic?
Shopify is the quickest way to put a store behind your name — try it free.
Four things that remove most of the difficulty
Talk to a person, not a lens — have someone stand behind the camera and explain it to them. This single change fixes more delivery problems than any coaching. Never memorise a script; use four bullet points and speak, because reading is instantly visible.
Start mid-sentence. No greeting, no name, no throat-clearing — begin with the explanation and add the introduction later if it is needed at all. Record in one block, without watching anything back until the session is over.
The things you hate are not what viewers see
Doctors watching themselves back fixate on their voice, their accent, a hand movement, the way they say one word. Patients notice none of that — they notice whether the explanation was clear and whether the person seemed trustworthy. The gap between what you hear and what they hear is enormous, and it never fully closes.
Which is why the decision to publish should not be yours alone in month one. Let someone else pick the take. Doctors reliably choose the worst one, because they choose the one where they made the fewest mistakes rather than the one where they sounded most like themselves.