Research

When a speech coach should say: I can't tell

Why useful speech coaching needs abstention: some recordings support playback or timing but not confidence, clarity, energy or causal advice.

A speech coach that always has an opinion is not automatically a better coach.

Sometimes the recording is incomplete.

Sometimes the microphone changed.

Sometimes the metric is descriptive but has no validated “better” direction.

Sometimes the listener question cannot be answered from acoustics.

Sometimes the system is being asked to infer confidence, authority or personality from a voice signal that does not support that conclusion.

In those cases, the useful answer is not a softer score.

It is:

I can't tell from this evidence.

Abstention sounds like a product weakness because AI products are trained to be helpful on every screen.

For speech coaching, it can be a core feature.

There is no public Ptichi desktop release today. This article explains the design principle we are building toward and gives you a manual version you can use with an ordinary recorder.

The dangerous path from signal to story

A voice recording contains measurable patterns.

Software may be able to estimate:

  • duration;
  • silent regions;
  • fundamental frequency descriptors;
  • recorded energy;
  • spectral features;
  • recognized words;
  • timing relationships.

The problem begins one layer later.

A pitch descriptor becomes:

You sound uncertain.

A louder second take becomes:

You sounded more confident.

A slower recording becomes:

Your clarity improved.

A long silence becomes:

You hesitated because you were nervous.

These statements are not merely friendlier labels for the measurements.

They are new claims.

Each requires evidence that the underlying signal supports the interpretation.

A speech coach can be technically correct about the waveform and still invent a story about the person.

One recording can be usable for one question and useless for another

This is the simplest model Ptichi uses internally.

Imagine a recording with a known gain change halfway through.

You may still be able to:

  • replay the words;
  • inspect where the sentence began and ended;
  • check whether the final condition was present.

You may no longer be able to make a clean claim that the speaker became louder.

The audio did not become “bad” in one global sense.

One capability became less trustworthy.

Now imagine a short recording with clean audio but too little material for a stable pitch summary.

Playback can be fine.

Word recovery can be fine.

A pitch-based interpretation may be unavailable.

This is why Ptichi's design avoids one microphone-quality score or one analysis-confidence percentage.

The useful question is:

Trusted for what?

Capture can change the metric before the speaker changes

The recording chain matters.

Awan and colleagues compared prerecorded speech and vowel samples from 24 speakers using several smartphones and a reference system. Device, environment and mouth-to-microphone distance affected selected acoustic measures differently, while strong relationships remained for other measures in the studied setup. Study abstract and publication details.

That is exactly the kind of result that should make a coach more specific.

Not:

Phones are inaccurate.

And not:

The microphone does not matter.

Instead:

This type of comparison may be sensitive to these capture conditions.

Suppose Take B has higher recorded energy.

Did the speaker project more?

Did the microphone move closer?

Did automatic gain change?

Did the room change?

Did the app use a different processing mode?

Until the comparison controls enough of those alternatives, “more energy” is a signal description, not a coaching conclusion.

This is the reason behind our recording-comparability longread.

Missing data must not become a speaking behavior

A second failure mode is more subtle.

Software often prefers continuity.

Graphs look better when lines connect.

Dashboards look better when every field contains a number.

But missing evidence is not zero.

If a live observer misses a section of audio, drawing a continuous line through the gap creates invented data.

If analysis has no supported observation for 600 milliseconds, calling that interval “silence” creates a speaking claim from a software absence.

If speech recognition cannot recover a phrase, calling it a pronunciation error may confuse recognition failure with speaker behavior.

Our “When a voice graph goes blank” article covers this failure boundary in more detail.

The principle is simple:

Unknown is a state. It is not an invitation to guess.

A number can be real and still have no coaching direction

Not all abstention comes from bad data.

Sometimes the metric is valid and the interpretation is not.

Speaking rate is the clearest example.

Suppose the system estimates 158 words per minute accurately.

What should it say?

“Too fast”?

“Good”?

“Slow down”?

The number alone cannot answer those questions.

Rate research shows why. Different listening tasks and populations respond differently to changes in speed. Both very fast and very slow speech can create costs under certain conditions. A descriptive pace measure does not contain its own universal target.

The same problem appears with pitch range.

A larger range is not automatically better.

A narrower range is not automatically monotone in a harmful sense.

Prominence depends on context and meaning, and research on prosody describes many-to-many relationships between acoustic form and information structure. Prosody and Information Structure.

So a coach may be able to say:

Your pitch range changed between these two comparable recordings.

while being unable to say:

The second range is better.

That is not incomplete analysis.

It is a correctly bounded claim.

Confidence is not a waveform label

The pressure to over-interpret becomes strongest around attractive human traits:

  • confidence;
  • authority;
  • charisma;
  • warmth;
  • leadership;
  • competence.

People absolutely make social judgments from voices.

That does not mean a consumer recording gives a coach a stable ground-truth confidence variable.

Research on public-speaking anxiety provides a useful caution. Self-report, observed behavior and physiological responses are related but not interchangeable. People can also judge their own performance differently from external observers. Measuring Public Speaking Anxiety: Self-report, behavioral, and physiological.

That study does not answer whether AI can estimate confidence.

It demonstrates the deeper problem: even a construct such as anxiety does not collapse neatly into one observable channel.

A voice coach should therefore be extremely cautious before turning acoustic proxies into personality conclusions.

Ptichi's rule is stricter:

Describe supported speaking behavior. Do not infer hidden personal traits because the interface wants a score.

Listener outcomes are different from acoustic outcomes

Suppose the goal is:

Make the condition easier to understand.

The speaker changes pitch, timing and articulation.

The acoustic features move.

Did the condition become easier to understand?

You still need a listener-relevant test.

An acoustic change can be evidence that something changed.

It does not automatically establish that the intended information became easier to recover.

This distinction appears repeatedly across Ptichi's content:

The listener task is the construct.

The acoustic metric is supporting evidence.

When the supporting evidence cannot answer the construct, the coach should not pretend it can.

Even human ratings need boundaries

“Use humans instead of AI” does not solve the problem automatically.

Human listeners can disagree.

A listener who knows the script is different from a listener who hears the message cold.

A colleague familiar with your voice is different from a first-time listener.

A trained evaluator is different from the actual audience.

Pairwise judgments are useful because they can turn a vague question into a concrete comparison. Current voice-evaluation products such as Voice Arena use blind pairwise human judgments and explicitly scope what their naturalness leaderboard measures. Voice Arena methodology.

That is a market-methodology example, not evidence for Ptichi.

The useful lesson is that even a human judgment needs a defined question.

Which version makes the condition easier to recover?

is stronger than:

Which voice is better?

Abstention applies to humans too.

If two takes are effectively tied, “no meaningful difference” is a valid answer.

The “cannot tell” ladder

A practical speech coach can have several levels of abstention.

1. Capture cannot support the measurement

Example:

Input gain changed between Take A and Take B. Energy comparison is unavailable.

The recordings may still be replayable.

2. Measurement exists, but direction is not validated

Example:

Take B had a wider pitch range. Ptichi cannot conclude that wider was better for this task.

3. The measurement does not answer the listener question

Example:

Pause duration changed, but we cannot infer that the condition became clearer from timing alone.

4. The construct is outside the product boundary

Example:

This recording cannot establish confidence, competence or personality.

5. There is too little evidence

Example:

The sample is too short for this check.

6. The evidence conflicts

Example:

The target metric moved in the intended direction, but listener recovery did not improve and effort increased.

A mature coach should handle all six without collapsing them into one low-confidence badge.

Try an abstention test with your own recorder

Choose this message:

We can release on Friday if security approves the exception by Thursday. Otherwise the release moves to Monday.

Record Take A.

Now create four deliberately awkward comparison cases.

Case 1 — move the microphone

For Take B, lean much closer while trying to “sound stronger.”

Question:

Can you confidently attribute a louder waveform to stronger delivery?

Correct answer: not from this pair alone.

Case 2 — change too many things

For Take C, slow down, rewrite the message, emphasize “Friday” and move the condition.

Question:

Which change caused any improvement?

Correct answer: the pair does not isolate one cause.

Case 3 — keep the setup, change one boundary

Return to the original setup. Change only the boundary before the condition.

Question:

Can you compare whether the condition is easier to recover?

Much better.

You can still avoid claiming durable learning.

Case 4 — ask for a trait

Listen to the final recording and ask:

How confident am I from 0 to 100?

The useful answer is not hiding in the waveform.

Replace it with:

Does the release sound confirmed or conditional?

Now the recording has a job it can actually support.

What abstention should feel like in a product

Bad abstention is:

Error.

or:

Confidence 42%.

Good abstention explains three things:

  1. What cannot be concluded.
  2. Why.
  3. What the person can still do.

For example:

Ptichi cannot compare recorded energy because the input level changed. Both recordings are still available. Replay them for the condition boundary, or repeat Take B with the same setup.

That is not an empty state.

It is a routing decision.

A useful system preserves the evidence that remains valid and suppresses the claim that became unsafe.

Why product teams resist this

Abstention creates uncomfortable screenshots.

A competitor can show:

Confidence: 84

while your product shows:

Cannot tell from this recording.

The first looks more intelligent.

The second may actually contain more reasoning.

Product teams also fear that unavailable states feel broken.

That is a legitimate UX problem.

The answer is not to invent results.

The answer is to make the boundary useful.

“Cannot tell” should lead to the smallest recovery action:

  • repeat with stable capture;
  • provide a longer sample;
  • ask a listener question;
  • use self-listening;
  • choose a supported target;
  • continue without the unavailable metric.

If no recovery action would make the claim trustworthy, the product should say that too.

Our read

AI speech coaching should be judged partly by the quality of its refusals.

That sounds unusual because benchmarks usually reward correct answers, not correct silence.

But a coach can fail in two directions:

  • miss a useful problem;
  • invent a problem that is not supported.

The second failure is especially dangerous in voice coaching because people can become self-conscious about normal variation.

A system that tells someone to lower their pitch, remove an accent feature or sound more authoritative without a listener-relevant reason is not merely noisy.

It may create the problem it claims to solve.

Ptichi's planned Speech Coach Benchmark therefore treats false coaching and abstention as first-class dimensions rather than footnotes.

The ideal coach does not maximize the number of observations.

It maximizes the number of supported, useful decisions.

What would change our view

We would narrow this position if strong validation showed that a particular metric supported a reliable coaching direction across the exact conditions where we currently abstain.

For example:

  • a capture perturbation may turn out not to matter for a specific timing metric;
  • a listener study may validate a particular boundary descriptor;
  • a language-specific analysis may become robust enough to support a clearer conclusion.

Then the system should become more informative.

Abstention is not permanent pessimism.

It is a temporary boundary around what has earned a claim.

The opposite is also true.

If validation reveals a metric is fragile, the product should become less confident.

That is how an evidence-based system should evolve.

A benchmark should punish false certainty

This gives us a different way to evaluate speech coaches.

Do not ask only:

Did the model identify the target?

Also ask:

Did it know when not to coach?

Useful benchmark cases include:

  • clipped input;
  • changed gain;
  • unsupported language;
  • too little material;
  • harmless accent variation;
  • a neutral recording with no material problem;
  • metric change caused by capture conditions;
  • a request to infer leadership or personality;
  • an intervention that improves one descriptor while making the delivery less natural.

A coach that gives eloquent advice in every case should score badly.

That is one of the most important ideas behind Ptichi's editorial/evidence policy and Method Card work.

What this does not prove

Ptichi has not publicly validated a general abstention engine.

The desktop product is still in development.

The external studies cited here support specific boundaries around capture, construct interpretation, naturalness and multimodal judgment. They do not prove Ptichi's exact rules.

The manual examples are authored reasoning exercises.

They are meant to make the evidence problem visible.

Continue with a question the evidence can answer

If Take A and Take B may not be comparable, start with Can you really compare two voice recordings?.

If the graph or automatic observation disappears, read When a voice graph goes blank.

If the recording itself sounds unfamiliar, use Why your voice sounds different on a recording.

The broader Ptichi technology direction is described in How Ptichi works.

The common rule is simple:

A speech coach should be precise about what it knows — and equally precise about what it does not.

What you can verify on this page

  • This page includes a Ptichi-authored example built for explanation or rehearsal.

    What it does not support: It is an editorial example, not an observed-user result, experiment or proof that Ptichi improves speech.

  • This page includes a bounded listen-and-compare exercise that you can run with your own recorder.

    What it does not support: The exercise does not prove that Ptichi improves speech or that a second take will generalize to other listeners or situations.

Sources and boundaries

  1. Smartphone Recordings are Comparable to “Gold Standard” Recordings for Acoustic Measurements of Voice primary-research-abstract · 2026-10-02

    What it supports: In controlled re-recordings of speech and vowels from 24 speakers using four smartphones and a reference system, device, setting and distance affected selected acoustic measures differently. Strong cross-device relationships and regression results supported comparability under the studied conditions.

    What it does not support: Abstract-level review only, using the public institutional abstract and bibliographic record. Restricted to the studied devices, conditions and measures. Does not validate Ptichi, its live graph or thresholds, speech-practice effectiveness, clinical use or a universal equipment recommendation. Used as context, not evidence that the study caused Ptichi’s architecture decisions. Journal issue: July 2025; the source ID is not a publication-date claim.

  2. Measuring Public Speaking Anxiety: Self-report, behavioral, and physiological primary-research · 2026-09-05

    What it supports: Self-report, observer ratings, behavior and physiological reactivity are not interchangeable during a speech challenge; socially anxious participants may underrate their performance relative to observers.

    What it does not support: Does not validate diagnosing anxiety from a voice recording or prove a Ptichi pressure-rehearsal intervention.

  3. Perceived Naturalness and Acoustic Characteristics of Habitual Versus Clear Speech for Individuals With Hearing Loss primary-research · 2026-09-27

    What it supports: Twelve adult female talkers read sentences in habitual, untrained clear, and trained clear conditions; 28 normal-hearing adults rated naturalness. Trained clear speech was rated least natural, was slowest, and had the most and longest pauses, with substantial variation between talkers and listeners.

    What it does not support: The study rated naturalness and measured acoustics, not listener intelligibility, real-world transfer, an ideal rate or pause length, or Ptichi training outcomes. The reading task and samples limit generalization.

  4. Prosody and Information Structure review · 2026-09-17

    What it supports: Reviews qualified evidence that speakers use prosody in relation to focus/givenness and listeners attend to prosodic cues when processing information structure; production shows many-to-many form-meaning mappings.

    What it does not support: Does not define one acoustic recipe for prominence, validate Ptichi scoring, or justify treating larger pitch/loudness changes as better communication.

  5. Voice Arena TTS Leaderboard — Methodology market-methodology-signal · 2026-09-14

    What it supports: Shows a current voice-AI evaluation product using blind pairwise human judgments, language-gated raters, issue tags and explicit limits on what its naturalness leaderboard measures.

    What it does not support: A TTS benchmark is not evidence about coaching human speakers, workplace comprehension, or Ptichi efficacy. It is included only as a market-methodology signal that human pairwise evaluation is actively used in voice products.

Reviewed: