Ptichi is being built as a desktop system for deliberate speech practice. Not as a test that reduces a person to one score, and not as an endless AI conversation that comments on every sentence. The central object in Ptichi is a controlled change between two attempts.
How Ptichi works: the speech-practice technology we are building
Inside Ptichi: programmes, prepared practice text, two-take comparison, local recording, bounded feedback, retests and a history of the skill.
You choose a concrete task, speak, listen to the first attempt, change one thing, speak again and compare. Then you try the same skill on different wording and, later, retest it.
In the shortest form:
task → Take A → listen → one cue → Take B → changed wording → later retest
The outside should feel simple. Under that simple loop we are building several technical layers so that the comparison is actually worth trusting.
The current status matters. Ptichi.site already provides free exercises and text-based ways to practise. The desktop app for macOS on Apple Silicon and Windows 11 x64 is in development; there are no verified public installers yet. Speech following and some automatic measurements remain planned or experimental directions, not released capabilities.
Why analysis alone is not the product
Speech can easily become a dashboard.
Software can measure pause duration, speaking rate, energy, pitch movement, word counts and many other signals. The hard question comes after the measurement: what should the person do next?
A precise number may still answer the wrong question.
A rate of 153 words per minute does not tell you whether that pace served this explanation. A pitch trace does not automatically tell you whether the listener heard the intended contrast. A model that outputs “confidence 82” has not turned confidence into a directly measurable property of the voice.
Ptichi therefore starts with a practice task rather than a metric.
Examples include keeping an important condition audible at the end of a sentence, making the boundary between two ideas clear, using silence to plan instead of spiralling into fillers, or explaining a technical point so the listener can recover its structure.
A measurement is useful only when it helps answer a bounded question like that.
The core mechanism: the Listening Loop
The working name for the central practice cycle is the Listening Loop.
Instead of “record something and receive a report”, the product should move through a deliberate sequence.
Choose what to practise
Starting from an empty Record button is a poor default. If someone came to work on interviews, explanations, pauses or statement endings, Ptichi should offer a meaningful task immediately.
That is why the product is organised around programmes and a practice library.
A programme is more than a folder of exercises. It defines what is being trained now, why it matters to the listener, what counts as a useful change and what the next step should be.
Use prepared practice material
We call the prepared material a practice score.
It is not a script with one correct performance. A score can mark a thought boundary, a word carrying the contrast, a condition that must not disappear, or a place where the ending deserves attention.
For example:
We can start on Friday / but only after the safety check.
The slash does not mean “pause for exactly 430 milliseconds”. It marks the practice question: can the listener still hear that the second part limits the first?
Prepared material gives us something important: reproducibility. We know which text version was used, which skill was targeted and which cue belonged to that attempt. That is more useful for training than inventing a new instruction during every recording and then trying to work out what changed.
Verify the recording before trusting analysis
A stream starting successfully does not prove that the recording is usable.
The selected microphone may produce a weak signal. A USB device can disappear mid-take. The operating system or driver may apply processing. A microphone mode can change between attempts. A development build can behave differently after packaging.
Ptichi therefore has a separate reliability concept: the Microphone Passport.
Its job is not to rank microphones. It should answer narrower questions: is there usable speech signal, which device and capture conditions were used, did something important change between takes, can a particular measurement be trusted, and should analysis be suppressed until the recording is repeated?
If the capture context changes, the system should not pretend that Take A and Take B are perfectly comparable.
Take A should be heard by the person first
We do not want AI to become the first voice in the room.
After Take A, the user should be able to listen and understand the task before the main correction appears. That is part of the training method: the person learns to notice evidence rather than only waiting for a model to grade it.
Automation can help next, but the response should be bounded.
Not:
Here are 17 things to fix.
More like:
Check the ending of the key statement. Make another take while keeping the rest roughly the same.
The working object for that feedback is an Intervention Card: one correction or listening question, the reason it matters, the limit of the advice and a transfer task.
Why one change is often better than twenty
Speech is multidimensional. Breathing, articulation, rate, pauses, loudness, pitch, rhythm, sentence structure, word choice and emotional delivery can all move at the same time.
If everything changes at once, the comparison collapses.
Ptichi tries to keep a small experiment:
What happens if we change only X?
Human speech is not really a set of independent sliders. But training becomes easier to inspect when the number of changing variables is reduced as far as the task reasonably allows.
Take B is therefore not a second chance to “sound better”. It is a test of a specific intervention.
Take B does not prove learning
A strong second attempt may simply mean that the person remembered the cue or copied the original wording.
That is why A/B comparison should be followed by transfer: the same skill on different words.
If the task is a thought boundary, change the sentence. If it is an interview answer, change the question or the example. If it is a statement ending, apply the cue to another statement.
Later comes a separate layer: retest. The skill is checked again after time has passed rather than only inside one successful session.
Progress is not one line on a chart
A training product can become misleading when it compresses progress into a single number.
Ptichi is designed to keep several states separate:
- Completion — the exercise was actually done.
- Controlled change — the intended change appeared in the rehearsed attempt.
- Transfer — it survived different wording or context.
- Retention — it appeared again in a later retest.
These are different claims.
Finishing lessons does not prove that speech changed. A good Take B does not prove transfer. Transfer in one session does not prove retention.
The planned Learning State therefore records what was practised, what could be reproduced, what is still unknown and what should be trained next instead of pretending that one universal speech score explains everything.
Automatic analysis should be modular, not omniscient
Different checks deserve different levels of trust.
One task may need only boundary timing. Another may use ASR to follow a prepared text. Another may compare an energy contrast. In some cases an automatic conclusion may simply not be reliable enough, and playback plus a listening question is the better product.
This leads to a capability-specific trust model. The question is not “do we trust the Ptichi analyser?” but “do we trust this capability, in this language, with this recording condition, for this bounded claim?”
If the system cannot support a correction, the right behaviour is to say that the automatic check is limited, not to invent a confident answer.
Real-time help should stay quiet
We are exploring a real-time layer, but its role is deliberately narrow.
During a take, the system should first preserve the recording, show capture state, optionally follow prepared text and recover from repeated words, skipped lines or moving backwards. If it loses its place, it should stop guessing and let the person recover manually.
We do not want the speaker to talk while watching a stream of scores and trying to satisfy five indicators.
Live support should therefore be quiet. The main coaching result comes after the attempt.
Why local-first matters
Speech practice often contains work details, names, unfinished thinking, interview answers and material that was never meant to become a cloud object.
The desktop product is being designed so that recordings and core analysis stay on the device by default.
That affects the product beyond privacy. Local history becomes a normal part of practice. Replaying and comparing attempts can be fast. Diagnostics can remain metadata-first instead of attaching speech or transcripts by default. Capability choices can also take the actual computer into account.
The current Ptichi website itself does not record audio or request microphone access. Until the desktop app is publicly released, recording exercises use a recorder the person already trusts.
Reliability is part of the technology
A speech product can look polished long before it can distinguish its own failure from the speaker’s failure.
We therefore separate four questions:
Does the software behave correctly?
Does a measurement agree with independently reviewed audio?
Does it run reliably on the real target computer and device?
Does the practice help a person?
Passing one does not pass the others.
A unit test does not prove measurement quality. A good algorithm result on a prepared file does not prove that a USB microphone survives suspend and resume. A reliable recording does not prove that the exercise improves a human skill.
Some conditions can be tested deterministically. Others require physical Windows/macOS hardware and packaged-build evidence. When that evidence does not exist yet, the product should say so.
What the user should finally see
The system can be deep inside without forcing the user to operate the depth.
The first-layer product model is intentionally small:
Home — what should I train next?
Train — programmes, library and the current practice.
Recordings — attempts, playback and comparison.
Progress — what changed, what is still unknown and what to work on next.
Settings — utility, not the centre of the product.
Measurement provenance, confidence boundaries and technical diagnostics should be available when needed, but they should not become the first screen.
What exists now and what we are building
Available on Ptichi.site now
- free exercises and guides;
- text practice without recording on the website;
- self-listening material;
- an English text helper for preparing a message;
- published practice transcripts and technique explanations.
In development
- a desktop app for macOS Apple Silicon and Windows 11 x64;
- local recording and playback;
- Take A / Take B comparison;
- programmes and a practice library;
- recording history and learning state;
- Microphone Passport and clearer capture-readiness checks.
Planned or experimental directions
- robust following of prepared text while a person speaks;
- bounded automatic observations for specific skills;
- transfer and later retests as first-class product steps;
- richer authoring for practice material;
- longer-term history of specific speaking skills.
These are development directions, not claims that the capabilities are already publicly available.
Where AI fits
AI is a component, not the product thesis.
It can help recognise speech, align an attempt with prepared text, locate a relevant segment, draft markup or provide a bounded semantic observation.
But the product question stays the same:
What should the person change in the next attempt, and how will we know that this particular thing changed?
If AI cannot answer that reliably, adding AI does not make the practice better.
The durable strength of Ptichi should come from the combination:
programme + prepared material + reliable capture + bounded feedback + comparison + transfer + history
Models will change. A useful practice protocol and the user’s own longitudinal record should survive model changes.
What kind of project Ptichi is meant to become
Ptichi is not ultimately a voice recorder.
We are building a workbench for deliberate speech practice.
A person may arrive with an interview, an assessment, a presentation, a technical explanation, a story from experience, a difficult question, second-language speech, microphone work or a general wish to sound clearer.
Instead of “improve your voice”, the system should turn that goal into a sequence of small, inspectable practice problems.
The end state is deliberately a little paradoxical: the better the training works, the less the person should need Ptichi.
First the product helps you notice a change. Then reproduce it. Then transfer it to a harder context, remove support and retest it later.
The goal is not to become good at using Ptichi. The goal is for the speaking skill to remain when Ptichi is closed.
Try the principle before the app is released
You can already run the simplest Ptichi loop with any recorder.
- Record Take A.
- Choose one thing only: a pause, a statement ending, a contrast or a thought boundary.
- Record Take B without deliberately fixing everything else.
- Compare only the chosen feature.
- Say a new sentence using the same technique.
If the change disappears on step five, that is useful information. You did not fail a test. You found the current boundary of the skill.
That boundary is exactly what Ptichi is being built to train.
