Research

When a voice graph goes blank

Why we are building recording and feedback as separate parts of Ptichi.

A missing live observation does not tell you that you stopped speaking. It does not tell you that your recording is safe, either. It tells you that the graph cannot provide a supported observation for that moment.

Development source review · 2 October 2026: This article describes the desktop code and its limits. No native build, runtime check or microphone test was run for this article. There is no public desktop release or installer. The website does not record audio or request microphone access.

That small distinction has become an important design rule in Ptichi: protect the recording, and be clear about what each layer of feedback actually knows.

Imagine rehearsing a short explanation before a call. You want the condition at the end to remain clear: “We can send the draft today, provided the final figures arrive.” During the second attempt, the live graph freezes and then disappears. Should you speak louder, repeat the sentence or stop to check the recording?

This is an illustrative design scenario, not a reported user incident. It shows why a practice app needs to distinguish a recording problem from a feedback problem. Otherwise, the person may try to correct their speech when the uncertainty belongs to the software.

The recording and the graph have different jobs

During a recording, the microphone delivers small blocks of digital audio. Our audio library, CPAL, provides these blocks through a callback: a function the audio system calls as data arrives. Processing that input and drawing a graph are separate operations. CPAL 0.18.2 documentation.

In the current development code, the main capture path first accepts valid samples into a prepared memory buffer. A separate, limited queue supplies the live display with observations of the input. If that queue becomes full, it can drop live observation blocks without removing the samples already accepted by the main path.

Source audio here means the samples the app received. The operating system may already have processed them; preserving these samples does not prove an unprocessed microphone signal.

This gives the two paths different priorities. The recording needs to preserve what arrived. The live display needs to tell us what is happening recently enough to be useful.

An old observation can be accurate about an earlier moment and still be misleading now. Suppose the graph shows a healthy level just as the observation worker stops responding. Leaving that level on screen would make old information look current. Replaying a backlog of observations as if they had just happened would cause the same problem.

The October 2 source changes address both cases. Old work is rejected, and live values expire when their supporting data is no longer fresh. The current native freshness limit is 300 milliseconds. It is an engineering rule for this display, not a recommended pause length or a threshold for good speech.

The trade-off is visible: the graph may provide less information. We prefer that to presenting a value whose time no longer matches the speaking moment.

There is a limit to this separation. Both paths still use the same computer, and copying data for the live view takes processing time. Code structure alone does not prove that every microphone and operating system will keep up. That needs physical testing.

A missing observation is not a measured pause

Three different kinds of time are involved.

The recording has a timeline based on the audio samples it accepted. The native observation system tracks how recently it received and published information. The interface has its own refresh and animation timing.

These clocks answer different questions. The interface waiting longer for a response does not mean that it measured more audio.

Consider a simplified example. A live view receives an observation covering 0.0–0.1 seconds, then another covering 0.4–0.5 seconds. It has no live observation covering the interval between them. Drawing a continuous line across that gap would be a visual guess. Counting the gap as silence would turn missing information into a claim about the speaker.

Our current source keeps the actual observation windows and rejects overlapping windows rather than counting the same audio twice. It also holds back summaries when the available windows do not provide enough coverage.

This matters even when nothing crashes. The live view keeps the latest observation, and the interface can miss an intermediate update. A useful live view is therefore not automatically a complete record of the take.

For a complete post-take observation, the retained audio is the appropriate source. A separate development function already examines every sample of a supported short WAV file. Connecting that foundation to a complete post-take Voice Trail remains further work. It is not a released import or analysis feature.

The reader consequence is simple: a blank section needs an explanation of the missing information. It must not become “you hesitated here” unless the system has the evidence required for that claim.

“Recording”, “saved” and “analysed” must remain separate

Keeping the graph separate does not mean recording without permission, continuing after Stop, or saving every sound the microphone receives.

The desktop workflow requires microphone access and an explicit recording action. Opening Home can check microphone availability, but does not request access or start a recording. A microphone check is also different from a practice take: checking the input does not make it a saved practice attempt. The live graph itself neither starts a new take nor gives permission to retain one.

Once an authorised practice recording is stopped, saving it has several steps. The source must be valid. The audio file must be written. The app must associate it with the correct attempt. Only then should optional analysis add its result.

This order is an implemented part of the current storage work. In the lesson path, the source and its exact attempt association are written to storage before optional analysis. If that analysis returns an error or fails, the saved recording can remain available with the analysis marked unavailable.

That is different from declaring every interrupted recording successful. Invalid input must still be reported. A saved valid prefix of an interrupted take does not prove that the whole intended take was captured. A failed save must not quietly attach an older recording to the new attempt.

There is also a clear recovery limit. The active source buffer remains in memory until Stop. A process crash before the source is written can lose the in-flight take. Recovery of an already written, pending save does not remove that earlier risk.

Our architecture decision to retain source audio and version its analysis helps explain the priority. An analysis method can change later. A saved recording gives us something to inspect again. Retaining only the latest score would leave no reliable way to check what that score described.

It also costs storage and requires careful deletion, export and recovery controls. Retention is useful only when the person understands and controls it.

A trustworthy value can still answer the wrong question

Even complete audio does not make every interpretation valid.

A digital level describes the recorded signal. It is affected by the recording path and settings. It is not a direct measure of vocal effort or how clearly someone explained a condition. An estimate of pitch does not establish confidence. Sound above an activity threshold is not automatically speech.

Ptichi currently keeps those distinctions explicit. Live pitch is experimental. The activity threshold is not a validated detector of speech or linguistic pauses. Providing a reference sentence does not prove that the person spoke those words, so its word count cannot simply become a speaking-speed result.

Research also gives us a reason to avoid one blanket label for recording quality. In Smartphone Recordings are Comparable to “Gold Standard” Recordings for Acoustic Measurements of Voice, Awan and colleagues recorded prerecorded speech and vowel samples from 24 speakers using four smartphones and a reference system. Device, room and distance affected the selected acoustic measures differently. The recordings nevertheless preserved strong relationships across devices, and statistical models could account for differences under the studied conditions. Study abstract and publication details.

Our read is that the useful question is “suitable for which measurement, under which conditions?” The study does not validate Ptichi's live graph, its thresholds or a speaking exercise. It also does not justify a general rule that expensive equipment is necessary. We use it here as context for the decision, not as evidence that it caused our earlier architecture choices.

For Ptichi, this supports keeping the status of each measurement separate. Playback may remain useful while an automatic comparison is unavailable. One supported observation must not make unrelated conclusions appear supported too.

We would reconsider a specific restriction if tests showed that the relevant measurement remained reliable across the conditions it currently excludes. We would still need separate evidence before claiming that the measurement helps a person practise.

Why not stop whenever the graph fails?

Stopping every time optional feedback becomes unavailable would make a short rehearsal depend on the availability of optional feedback. The person might still have a useful recording to finish and hear.

At the other extreme, hiding a recording failure behind a working animation would be worse. A moving graph cannot guarantee a complete, saved take.

The distinction we implement is narrower: an observation failure must not automatically discard valid source audio. A capture failure remains a capture failure and needs its own visible state. Stop and discard controls must stay reachable while the app resolves what happened.

This is why the recent work includes tests for awkward timing: a delayed status response after Stop, a stalled observer, or a failed save followed by a retry. They describe software failure cases. They do not tell us how often users encounter them or whether a person understands the resulting message.

That second question is still open. “Unavailable” can be honest and unhelpful at the same time. The interface needs to explain whether the person can finish, replay, retry saving or make a new take. Observed use is needed to check that the distinction is clear without making practice feel like technical troubleshooting.

Local audio still needs a clear privacy boundary

The core desktop design keeps recordings, personal practice state and work text on the device by default. An unavailable measurement does not authorise a silent upload to another service for a second opinion.

Local storage is not the same as leaving no copies anywhere. A person may export a WAV, or the operating system may back up the application data. Deleting a managed recording cannot remove an exported copy or every backup.

The website is a separate boundary. It receives none of the audio from a recorder you choose for a web exercise. That recorder has its own storage and privacy settings. Our privacy explanation describes these distinctions.

A check you can use before any voice score

You can try the underlying principle with a recorder you already use. Choose non-confidential material and keep the setup comfortable and consistent. The existing microphone placement guide covers that setup question.

Record one short sentence with a condition at the end. Stop, confirm that the file exists, and replay the whole sentence before opening any automatic feedback. Can you hear the condition? Is the ending present? Did the recording include the part you intended to compare?

If the audio is incomplete, make another take. If the audio is usable but a metric is missing, leave that metric unknown. You can still listen for the specific feature you chose, without treating your judgement as proof of a learning effect.

For a complete speaking exercise, continue with Boundary contrast. It guides the move from one attempt to one change and different wording. There is no need to invent a score for the graph's blank space.

The development lesson is not that feedback has little value. Feedback has more value when the app can explain what supports it, what is missing and what the person can still do. A retained, valid recording gives practice somewhere to return to. A clear limit tells us when to listen instead of guessing.

What you can verify on this page

  • This page includes a Ptichi-authored example built for explanation or rehearsal.

    What it does not support: It is an editorial example, not an observed-user result, experiment or proof that Ptichi improves speech.

  • This page includes a bounded listen-and-compare exercise that you can run with your own recorder.

    What it does not support: The exercise does not prove that Ptichi improves speech or that a second take will generalize to other listeners or situations.

Sources and boundaries

  1. CPAL 0.18.2 — crate documentation primary-software-documentation · 2026-10-02

    What it supports: After an input stream starts, CPAL delivers captured audio samples through a periodically invoked data callback.

    What it does not support: Explains CPAL’s stream and callback model only. Does not establish Ptichi’s capture continuity, device reliability, latency, source preservation, saving, privacy, learning outcomes or release readiness.

  2. Smartphone Recordings are Comparable to “Gold Standard” Recordings for Acoustic Measurements of Voice primary-research-abstract · 2026-10-02

    What it supports: In controlled re-recordings of speech and vowels from 24 speakers using four smartphones and a reference system, device, setting and distance affected selected acoustic measures differently. Strong cross-device relationships and regression results supported comparability under the studied conditions.

    What it does not support: Abstract-level review only, using the public institutional abstract and bibliographic record. Restricted to the studied devices, conditions and measures. Does not validate Ptichi, its live graph or thresholds, speech-practice effectiveness, clinical use or a universal equipment recommendation. Used as context, not evidence that the study caused Ptichi’s architecture decisions. Journal issue: July 2025; the source ID is not a publication-date claim.

Reviewed: