Why I Built SpeechMe to Work Differently
Building SpeechMe became far more of an adventure, and at times more of a nightmare, than I had imagined.
In principle, it looked simple: build a smart layer over AI and use it to generate speeches. In reality, that was nowhere near enough.
I have built several apps, and each one has taught me something new about development and process. SpeechMe was no different.
Every Speech Needed More Data Than Expected
The first hit was realizing how much data needed collecting for each speech occasion. This wasn’t some trivial few questions, this turned out to be several weeks of in depth research into what a good, no, correct that, a great speech for every occasion consisted of.
Every speech occasion has its own unique set of ideal qualities and data sets, from a coding point of view anyway. This is why people attempting to write speeches with a generic AI fail, the prompt structure and the actual required content is so far removed from what the normal person realizes or can specify that the speeches rarely convey or even relate to what they really want.
Building the separate workflows, and finding the right questions for each one, was one of the hardest parts of the project.
The research became the foundation of SpeechMe: a guided speechwriting process at a price ordinary people can justify. The questions draw out the details a professional writer would need without expecting the user to write a professional brief or engineer an AI prompt.
The questions are split into stages, with each stage covering a different part of the speech. The app explains why the information matters and what kind of answer is most useful.
AI Text Is Not the Same as Human Speech
The next problem appeared as soon as the drafts were read aloud: correct grammar did not make them sound like natural speech.
A draft can look acceptable at first glance and still contain problems that become obvious aloud.
The biggest issue, especially with speeches, is that readable text is not the same as human speech. AI can produce a sentence that is technically correct, but technically correct is not enough. A speech has to be read out loud, in front of people, in a real situation. That is where a lot of AI-generated writing starts to fall apart.
There are lines that look fine on screen but sound completely off when spoken. There are phrases that feel too polished, too generic, too neat, or just not like something a real person would say in that moment.
A sentence can work on screen and still fail in the room. The audience hears pace, emphasis and tone as well as the words.
AI writing often repeats the same words, sentence shapes and transitions. It can overuse dashes and lists or produce language that looks polished but says very little.
Those patterns matter because listeners notice when a speech sounds manufactured, even if they cannot identify the exact phrase that caused the problem.
Along the same lines, another reason, particularly for speeches, is that when you actually read the speech out loud as if you were actually addressing your audience, you realize that a lot of things in there are just not what a human would actually say, they look ok, well sort of, but read aloud, they just don't sound right.
Real Human Speech Had to Be Built Into the Original Generation Process
This was not something that could just be fixed at the end. If the original speech generation was poor, then a later clean-up pass would only be patching a weak draft. The speech needed to be generated in a better way from the start.
The first step was to create strong examples for each part of each speech occasion. Those examples gave the generation process a clearer model of natural spoken language in context.
That work now sits inside the original speech generation process. The point was to make the first proper draft much stronger before any later refinement happened.
One early father-of-the-bride draft said, ‘I would like to thank everyone for being here today to celebrate Sam and Becky’. The grammar was fine, but the meaning was slightly wrong: the room was celebrating their marriage. The line needed to say exactly that.
It is a small correction, but speeches are full of details like this. Left unedited, they make the wording feel subtly wrong.
To make things just a little more difficult, as if there weren't enough things to do, another thing I had to account for was language style. A speech written for a UK wedding, an American rehearsal dinner or an Australian sports team event should not sound identical.
Supporting UK, US and Australian English meant adapting the flows, examples and wording for each locale rather than changing the currency and spelling at the end.
SpeechMe currently supports those three English variants, with room to add more later.
Training the AI on what is real human speech was actually quite a task and something I did not realize would be so demanding as we all assume that AI is trained in this at its base level. Unfortunately until I came to actually build this out, I didn’t realize how bad it actually is.
Why the Humanizer Had to Exist
Even after the main generation process improved, a separate clean-up stage was still needed.
Better questions and structures improved the first draft, but stiff or repetitive wording still needed a separate review. That led to the Humanizer.
I created a ‘Humanizer’, a process that takes the speech and runs it through a series of additional safeguards that check the generated speech afterwards and remove the remaining AI quirks, repeated phrasing, clichéd terms and wording issues that can make a speech sound off when read aloud.
The user can then review the refined draft and change any line that still does not sound like them.
Why One Generation Step Was Not Enough
Another issue to overcome was taking the actual data from the user answers and generating a great speech.
This, I discovered, requires more than one step.
I split the process into five stages so the user can review, edit or regenerate sections before moving on.
An initial outline followed by a draft, a punch up - no not a fight but a selection of targeted improvements, the Humanizer pass and finally the delivery. This 5 stage process ensures the user has full control of the generation with edits and regenerations and provides a speech built to replicate the kind of structure, refinement and control you would expect from a professional speechwriting process.
The aim is a speech that uses the user's details and still feels comfortable in their voice.
Guardrails for Material That Should Stay Out
There was a risk remaining though and that risk was not the AI, it was the user. For a lot of people, writing a speech, especially something like a best man speech where it is expected to be humorous, can lead to disaster.
The app therefore needed safeguards for embarrassing, inflammatory or plainly unsuitable material.
As a precaution, there are safeguards and settings that are specific to different speech occasions. Level sliders that allow you to set what is an appreciable level of humor and the tone of the speech with warnings and advice to help guide decisions.
Users can also specify topics, references or details that must not appear, even if they were mentioned elsewhere in the answers.
Building a Complete Solution
A good draft only solves half the problem. The speaker still has to stand up and deliver it.
I quickly realized that generating the speech was half the job, a great speech is no good if you fluff your lines. Further solutions were required and that's when the Rehearse and Teleprompter features fell into place.
SpeechMe now covers the writing process and the practical work of preparing to deliver the speech.
The Rehearse feature took time to get right. Instead of sending the written text straight to a voice, the app prepares a delivery version with pacing, pauses and timing instructions.
That structured version is then used to create rehearsal audio with pacing, pauses and delivery timing. It can be played in the app for you to listen to and speak along with, and the rehearsal mode lets you adjust the speed of delivery as you practice for the actual occasion.
The final issue was access to the speech on the day. A printed copy is available, but it is not the only option.
The mobile teleprompter offers a cleaner alternative for speakers who are comfortable using a phone.
Open SpeechMe on your phone and the finished speech is available in a full-screen teleprompter. Use automatic scrolling or move through the text manually at your own pace.
My aim was to cover the whole job: finding the material, shaping the speech, refining the wording and preparing to deliver it.
Why SpeechMe Uses a One-Off Price
While building SpeechMe, I compared professional speechwriting services with basic AI tools. A human writer offers interviews and bespoke revisions, but that service can be expensive. Basic tools cost less but leave the structure, humor, editing and rehearsal to the user.
SpeechMe sits between those options. The one-off price includes guided questions, an editable outline and draft, Humanizer checks, rehearsal audio, a teleprompter and PDF export.
The aim is not to replace a professional writer. It is to offer a guided alternative for people who want to work from their own material without spending hundreds of pounds.
I built the process to reduce the blank-page work while leaving the speaker in control of every line.
Now, Go SpeechMe.