spectral.rs

crates/veilvoice-core/src/spectral.rs

veilvoice-core · 442 lines · read the source here · or on GitHub

Frequency-domain de-identification transform.

For every STFT frame we:

  1. take the magnitude spectrum and discard the measured phase. This is the irreversible step, it permanently erases the speaker's waveform / micro-timing;
  2. estimate a smooth spectral envelope (the vocal-tract / formant structure, i.e. the biometric identity) and the excitation residual (glottal source + phonetic detail that carries the words);
  3. shift the excitation by a cryptographically-modulated pitch ratio and warp the envelope by an independent formant ratio, so the identity is moved somewhere it never was while the phonemes stay legible;
  4. resynthesise a fresh phase, plus a fixed random per-bin offset.

Voiced frames: an explicit harmonic comb

Step 4 has two modes. On unvoiced frames each bin accumulates its own centre frequency, which is the classic channel-vocoder phase, exactly right for fricatives and noise.

On voiced frames that alone is not enough, and it is audible. Bin centres are multiples of sample_rate / n (46.875 Hz at the default settings), and a harmonic peak spans several bins, so a plain channel vocoder turns each partial into a cluster of independent grid-frequency sinusoids with unrelated phases. A 210 Hz voice comes out beating around 187.5 and 234.4 Hz: metallic, and with a pitch that cannot be steered, which would make the canonical register crate::accent aims for unreachable.

So when the frame is voiced and accent neutralisation is active, the excitation is not resampled at all. It is replaced by an ideal harmonic comb at the canonical fundamental, quantised to the nearest whole bin. This is the textbook source-filter model of voiced speech (an impulse train through the vocal-tract filter), and because every comb line then sits exactly on a bin centre, the existing per-bin phase advance is precisely the right advance for it: successive frames overlap-add coherently and each harmonic emerges as one clean partial. The envelope still supplies the formants, so the vowels are untouched.

Snapping to the bin grid is what buys that coherence, and it costs pitch resolution, so the grid step is coarse. That is not a problem for the default configuration, which maps every speaker onto a single constant register that need only be snapped once; it does mean any residual intonation (prosody_flatten below 1.0) is quantised to the same grid. Lifting that restriction needs window-kernel synthesis, noted as future work in the project roadmap.

None of this weakens irreversibility. The measured phase is still discarded in full, and pinning the output to one canonical fundamental destroys more pitch information than randomising it would, because a constant carries nothing.

Between steps 2 and 4 the optional crate::accent neutraliser folds in its long-term corrections: it reads the unwarped envelope to measure the speaker's vocal-tract scale, contributes extra pitch and formant ratios, and rotates the warped envelope toward a canonical spectral tilt.

The measured phase is never reused, so no amount of downstream processing can reconstruct the original excitation phase: the transform is one-way.

In plain words

This is the part that actually destroys the voiceprint, and the step that cannot be undone.

Sound carries two things: which frequencies are present, and how they line up in time. The second one, the timing, is a great deal of what makes a voice recognisably yours, and it is thrown away here and replaced. It is not scrambled or hidden; it is discarded, and there is nothing left to recover it from.

What is kept is enough for the words to stay clear. That is the whole trade: the sentence survives, the speaker does not.

WHAT THIS FILE CONTAINS

442 lines defining 5 functions (3 public), 1 type and 0 constants. Everything below is read out of the source, so it cannot disagree with the code.

The types it owns.

  • struct SpectralState line 80 · Persistent per-instance state for the spectral transform.

What happens when it runs. These are the ways in: public, and nothing else in this file calls them, so they are what an outside caller reaches first.

  • SpectralState::new line 104 · n = FFT size, hop = analysis/synthesis hop, rand_phase = fixed per-bin phase offsets in radians (length n/2+1) drawn from the CSPRNG.
  • SpectralState::retarget_phase_offsets line 145 · Aim the per-bin phase offsets at fresh values.
  • SpectralState::transform line 159 · Rewrite spec (length n/2+1) in place, given the current modulation.
    reaches box_smooth, resample_linear

WHAT CALLS WHAT

SpectralState::new line 104 SpectralState:: retarget_phase_offsets line 145 SpectralState::transform line 159 resample_linear line 293 box_smooth line 314 entry: a way in: public, and nothing in this file calls it helper: private to this file dashed: a call that goes back up, or across a wrapped rank The functions this file defines, and the calls between them. An edge means the callee's name appears, called, inside the caller's body. This is a syntactic reading, not a type-resolved one.

The functions this file defines, and the calls between them. An edge means the callee's name appears, called, inside the caller's body. This is a syntactic reading, not a type-resolved one.

The same graph as Mermaid source
%%{init: {"theme":"base","themeVariables":{"background":"#1a1b26","primaryColor":"#1f2335","primaryTextColor":"#c0caf5","primaryBorderColor":"#7aa2f7","secondaryColor":"#16161e","tertiaryColor":"#16161e","lineColor":"#737aa2","textColor":"#c0caf5","mainBkg":"#1f2335","nodeBorder":"#7aa2f7","clusterBkg":"#16161e","clusterBorder":"#2f3549","fontFamily":"ui-monospace, SFMono-Regular, Consolas, monospace","fontSize":"14px"}}}%%
flowchart TD
    n_new(["SpectralState::new<br/>line 104"])
    n_retarget_phase_offsets(["SpectralState::<br/>retarget_phase_offsets<br/>line 145"])
    n_transform(["SpectralState::transform<br/>line 159"])
    n_resample_linear["resample_linear<br/>line 293"]
    n_box_smooth["box_smooth<br/>line 314"]
    n_transform --> n_box_smooth
    n_transform --> n_resample_linear
    click n_new href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-core/src/spectral.rs#L104" "open the source"
    click n_retarget_phase_offsets href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-core/src/spectral.rs#L145" "open the source"
    click n_transform href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-core/src/spectral.rs#L159" "open the source"
    click n_resample_linear href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-core/src/spectral.rs#L293" "open the source"
    click n_box_smooth href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-core/src/spectral.rs#L314" "open the source"
    classDef entry fill:#1f2335,stroke:#7aa2f7,color:#c0caf5
    class n_new,n_retarget_phase_offsets,n_transform entry
    classDef helper fill:#1f2335,stroke:#bb9af7,color:#c0caf5
    class n_resample_linear,n_box_smooth helper

This site loads no third-party script, so it cannot run Mermaid; the diagram above is the same nodes and edges drawn by the generator instead. GitHub renders the source below directly.

ITEMS

ItemLineDocumentation
SpectralState pub struct80Persistent per-instance state for the spectral transform.
SpectralState::new pub fn104n = FFT size, hop = analysis/synthesis hop, rand_phase = fixed per-bin phase offsets in radians (length n/2+1) drawn from the CSPRNG.
SpectralState::retarget_phase_offsets pub fn145Aim the per-bin phase offsets at fresh values.
SpectralState::transform pub fn159Rewrite spec (length n/2+1) in place, given the current modulation.
resample_linear fn293Linear resampling of a non-negative spectral function.
box_smooth pub(crate) fn314In-place-ish box smoother using a running sum.