veilvoice-core

veilvoice-core

Irreversible voice de-identification DSP engine: cryptographically-modulated pitch/formant scrambling with preserved intelligibility.

reference · the same page on GitHub

The security-critical heart of VeilVoice: an irreversible, cryptographically modulated voice de-identification engine.

What it guarantees (and what it deliberately does not)

VeilVoice destroys the biometric voiceprint, meaning fundamental pitch, formant structure, timbre, accent and micro-timing, so that neither software nor a human can re-identify the speaker or reconstruct the original waveform. It does not hide the words: intelligibility is preserved on purpose, because a scrambler you cannot understand or transcribe is useless. "Fill the whole spectrogram with white noise" and "stay transcribable" are mutually exclusive; see docs/WHITEPAPER.md for the full argument.

Accent

AccentConfig additionally maps every speaker onto one canonical pitch register, vocal-tract scale and long-term spectrum, so the melody and colour of an accent, along with two of the strongest biometric features there are, do not survive. What no signal-level transform can remove is the segmental side of an accent: which phonemes were actually produced. At that level the accent and the words are the same thing, and changing it means changing what was said. See AccentConfig for the full argument and the limit, which the whitepaper must state rather than overclaim.

Why it is one-way

Every STFT frame has its measured phase discarded and resynthesised from scratch (see spectral). The original excitation phase, which encodes the precise waveform and a speaker's micro-timing, is never stored and never reused, so no downstream process can recover it. On top of that, the pitch and formant shifts are driven every frame by a ChaCha20 CSPRNG (modulation) whose seed never leaves the process and is zeroized on drop, so there is not even a single fixed transform to invert.

Example

use veilvoice_core::{Deidentifier, DeidConfig};

let mut deid = Deidentifier::new(DeidConfig::default()).unwrap();
let input = vec![0.0f32; 4800];
let output = deid.process_vec(&input);
assert_eq!(output.len(), input.len());
// Live processing cost, e.g. for a latency read-out:
let _ms = deid.stats().last_block_ms();

In plain words

This is the part that actually changes the voice.

A recording goes in and a recording comes out. The words are the same and you can still understand every one of them; the voice is not yours any more, and there is no setting, no key and no clever program that turns it back. What made it recognisably you -- the pitch, the shape of your mouth and throat, the timing, the music of your accent -- is not hidden. It is thrown away, and everybody who goes through it comes out sounding like the same handful of people.

What it does not do is keep your words secret. It is not meant to: a voice nobody can understand would be no use to anyone. If what you said would identify you, this has not touched that.

HOW THE CRATE FITS TOGETHER

Every arrow is a crate:: or super:: path one module actually uses, read out of the source rather than drawn by hand.

The same graph as Mermaid source
%%{init: {"theme":"base","themeVariables":{"background":"#1a1b26","primaryColor":"#1f2335","primaryTextColor":"#c0caf5","primaryBorderColor":"#7aa2f7","secondaryColor":"#16161e","tertiaryColor":"#16161e","lineColor":"#737aa2","textColor":"#c0caf5","mainBkg":"#1f2335","nodeBorder":"#7aa2f7","clusterBkg":"#16161e","clusterBorder":"#2f3549","fontFamily":"ui-monospace, SFMono-Regular, Consolas, monospace","fontSize":"14px"}}}%%
flowchart TD
    n_lib(["lib.rs<br/>89 lines"])
    n_accent["accent.rs<br/>701 lines"]
    n_chain["chain.rs<br/>1741 lines"]
    n_effects["effects.rs<br/>245 lines"]
    n_modulation["modulation.rs<br/>323 lines"]
    n_pitch["pitch.rs<br/>286 lines"]
    n_spectral["spectral.rs<br/>442 lines"]
    n_stft["stft.rs<br/>264 lines"]
    n_voices["voices.rs<br/>872 lines"]
    n_window["window.rs<br/>101 lines"]
    n_accent --> n_pitch
    n_accent --> n_spectral
    n_chain --> n_accent
    n_chain --> n_effects
    n_chain --> n_modulation
    n_chain --> n_pitch
    n_chain --> n_spectral
    n_chain --> n_stft
    n_spectral --> n_accent
    n_spectral --> n_pitch
    n_stft --> n_window
    click n_lib href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-core/src/lib.rs" "open the source"
    click n_accent href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-core/src/accent.rs" "open the source"
    click n_chain href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-core/src/chain.rs" "open the source"
    click n_effects href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-core/src/effects.rs" "open the source"
    click n_modulation href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-core/src/modulation.rs" "open the source"
    click n_pitch href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-core/src/pitch.rs" "open the source"
    click n_spectral href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-core/src/spectral.rs" "open the source"
    click n_stft href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-core/src/stft.rs" "open the source"
    click n_voices href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-core/src/voices.rs" "open the source"
    click n_window href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-core/src/window.rs" "open the source"

This site loads no third-party script, so it cannot run Mermaid; the diagram above is the same nodes and edges drawn by the generator instead. GitHub renders the source below directly.

THE FILES

FileLinesWhat it is
accent.rs701Accent and speaker-trait neutralisation.
chain.rs1741The assembled de-identification chain and its live performance statistics.
effects.rs245Light time-domain effects applied after resynthesis.
lib.rs89The security-critical heart of VeilVoice: an irreversible, cryptographically modulated voice de-identification engine.
modulation.rs323Cryptographically-seeded modulation of the effect parameters.
pitch.rs286Monophonic fundamental-frequency tracker (decimated YIN).
spectral.rs442Frequency-domain de-identification transform.
stft.rs264Streaming short-time Fourier transform with overlap-add resynthesis.
voices.rs872Destination voices: several canonical registers instead of one.
window.rs101Analysis and synthesis windowing, and the one constant that keeps overlap-add honest.
spectrum_report.rs107Where do the output partials actually land?
veil_a_buffer.rs54no module documentation yet
hostile_audio.rs389The engine against input that is not well-behaved audio.