voices.rs

crates/veilvoice-core/src/voices.rs

veilvoice-core · 872 lines · read the source here · or on GitHub

Destination voices: several canonical registers instead of one.

What this is for

By default every speaker VeilVoice processes comes out as the same voice, with one pitch register, one vocal-tract scale and one long-term spectrum. That is the many-to-one mapping the whole project rests on, and for a single speaker it is exactly right.

It is wrong for a conversation. Two people veiled into one indistinguishable voice produce a recording nobody can follow: the words survive and the turn-taking does not, so a listener cannot tell a question from its answer. This module hands out a small table of distinct destination voices, so a recording with three people in it comes out with three voices in it.

The security property, and how it survives

The property that matters is that the output voice is a function of the slot, not of the speaker. Every input mapped onto slot 3 comes out as voice 3, whoever they were. The mapping is still many-to-one, with infinitely many inputs per output, so there is still no inverse to compute. What changes is the number of buckets, from one to at most ten.

This is why a voice is never derived from the speaker. Choosing a destination by measuring the input is the obvious implementation, and the one that would sound most natural. It would make the output voice a function of the input voice, which is precisely the linkage the project exists to destroy. Slots are assigned by turn order, and turn order is something the user supplies.

What a conversation leaks that a monologue does not

Stated plainly, because it is a real cost and nobody should have to work it out for themselves:

  • How many people were talking. Ten voices in the output means ten speakers in the input.
  • Who spoke when, and for how long. Turn-taking structure is preserved on purpose, because it is the thing that makes the result usable, and turn structure is information about a conversation. Overlaps, interruptions, the length of each answer and the rhythm of the exchange all survive.
  • Nothing about who they were. The voiceprint of each speaker is destroyed exactly as thoroughly as in single-speaker mode: the same phase discard, the same many-to-one normalisation onto the slot's canonical values.

A register has to land on a bin, and that is the whole constraint

The first version of this table picked five registers about 26 Hz apart on the reasoning that the just-noticeable difference for the fundamental of speech is around 8 Hz, so 26 Hz would be three times that. The reasoning was sound and the table was wrong, because it measured the wrong thing.

When accent neutralisation is on, crate::spectral does not resample the excitation. It replaces it with a harmonic comb at the canonical fundamental, quantised to the nearest whole FFT bin so that every comb line sits on a bin centre and the frames overlap-add coherently. The rendered fundamental is therefore not the number in the table. It is

round(target_f0 / bin_hz) * bin_hz,   bin_hz = sample_rate / frame_size

At the default 1024-point frame and 48 kHz that spacing is 46.875 Hz, so the five registers 105, 131, 157, 183 and 209 Hz render as 93.75, 140.625, 140.625, 187.5 and 187.5, which is three distinct pitches, not five. Two pairs of speakers would have shared a register with nothing in the interface saying so. It was found by measuring the fundamental of an actual rendered file, not by reading the code, and the tests below now measure the same thing the ear would.

So the registers are bin-exact by construction: each is a whole number of bins at the default configuration. Inside the range where a resynthesised voice stays intelligible, roughly 90 to 240 Hz, below which the comb has too few harmonics under the vowels and above which it stops being a speaking register, there are exactly four:

| bin | rendered | |---:|---:| | 2 | 93.75 Hz | | 3 | 140.625 Hz | | 4 | 187.5 Hz | | 5 | 234.375 Hz |

Ten voices, from four registers and three vocal tracts

The second axis is the canonical vocal-tract scale, which is a continuous warp and is not quantised, so it is free to take values the ear can separate: 620, 760 and 900 Hz, each about 22 % from its neighbour, all inside the range where the vowels stay natural.

Four registers times three tracts is twelve, and MAX_VOICES ships ten of them. Twenty, which was asked for, is not available: it would need either registers a single bin apart at a frame size four times longer, which quadruples the latency, or vocal tracts close enough to be heard as the same person on a different day.

If you change the frame size, check the table again

The registers are exact at the default configuration. A caller who changes crate::DeidConfig::frame_size or the sample rate moves the bin grid underneath them, and two registers can collide again. Voice::rendered_f0_hz reports what a given configuration will actually produce, and distinct_voices counts how many of the ten survive it, so a front end can say "this frame size gives you six distinguishable voices" rather than handing out ten labels for six sounds.

In plain words

The set of voices a recording can be turned into.

By default everyone comes out as the same one. That is on purpose: if every speaker sounds identical, there is nothing in the result that distinguishes one from another, and nothing to trace back.

When a recording has several people in it that becomes a problem, because a listener cannot follow who is who. So there is a small set of destination voices to hand out instead, chosen to be as far apart as the arithmetic allows, and a measured limit on how many of them can genuinely be told apart by ear.

WHAT THIS FILE CONTAINS

872 lines defining 11 functions (11 public), 1 type and 11 constants. Everything below is read out of the source, so it cannot disagree with the code.

The types it owns.

  • struct Voice line 134 · One destination voice: the canonical values every speaker in this slot is mapped onto.

What happens when it runs. These are the ways in: public, and nothing else in this file calls them, so they are what an outside caller reaches first.

  • Voice::applied_to line 156 · Apply this voice to an AccentConfig.
  • Voice::checked line 183 · Whether this voice is inside the range the engine can render usefully.
  • Voice::describe line 227 · A short label for an interface: "low register, narrow tract".
    reaches rendered_f0_hz, bin_hz
  • voice line 326 · The destination voice for slot index.
  • clear_voices line 412 · How many voices can be handed out before two of them are too alike.
    reaches all, separation
  • closest_pair line 432 · The closest pair among the first count voices, as a ratio.
    reaches all, separation
  • distinct_voices line 458 · How many of the ten are still distinguishable under config.
    reaches all

WHAT CALLS WHAT

Voice::applied_to line 156 Voice::rendered_f0_hz line 170 Voice::checked line 183 Voice::describe line 227 bin_hz line 270 voice line 326 all line 337 separation line 363 clear_voices line 412 closest_pair line 432 distinct_voices line 458 entry: a way in: public, and nothing in this file calls it api: public, and also used inside this file dashed: a call that goes back up, or across a wrapped rank The functions this file defines, and the calls between them. An edge means the callee's name appears, called, inside the caller's body. This is a syntactic reading, not a type-resolved one.

The functions this file defines, and the calls between them. An edge means the callee's name appears, called, inside the caller's body. This is a syntactic reading, not a type-resolved one.

The same graph as Mermaid source
%%{init: {"theme":"base","themeVariables":{"background":"#1a1b26","primaryColor":"#1f2335","primaryTextColor":"#c0caf5","primaryBorderColor":"#7aa2f7","secondaryColor":"#16161e","tertiaryColor":"#16161e","lineColor":"#737aa2","textColor":"#c0caf5","mainBkg":"#1f2335","nodeBorder":"#7aa2f7","clusterBkg":"#16161e","clusterBorder":"#2f3549","fontFamily":"ui-monospace, SFMono-Regular, Consolas, monospace","fontSize":"14px"}}}%%
flowchart TD
    n_applied_to(["Voice::applied_to<br/>line 156"])
    n_rendered_f0_hz["Voice::rendered_f0_hz<br/>line 170"]
    n_checked(["Voice::checked<br/>line 183"])
    n_describe(["Voice::describe<br/>line 227"])
    n_bin_hz["bin_hz<br/>line 270"]
    n_voice(["voice<br/>line 326"])
    n_all["all<br/>line 337"]
    n_separation["separation<br/>line 363"]
    n_clear_voices(["clear_voices<br/>line 412"])
    n_closest_pair(["closest_pair<br/>line 432"])
    n_distinct_voices(["distinct_voices<br/>line 458"])
    n_clear_voices --> n_all
    n_clear_voices --> n_separation
    n_closest_pair --> n_all
    n_closest_pair --> n_separation
    n_describe --> n_rendered_f0_hz
    n_distinct_voices --> n_all
    n_rendered_f0_hz --> n_bin_hz
    click n_applied_to href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-core/src/voices.rs#L156" "open the source"
    click n_rendered_f0_hz href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-core/src/voices.rs#L170" "open the source"
    click n_checked href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-core/src/voices.rs#L183" "open the source"
    click n_describe href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-core/src/voices.rs#L227" "open the source"
    click n_bin_hz href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-core/src/voices.rs#L270" "open the source"
    click n_voice href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-core/src/voices.rs#L326" "open the source"
    click n_all href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-core/src/voices.rs#L337" "open the source"
    click n_separation href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-core/src/voices.rs#L363" "open the source"
    click n_clear_voices href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-core/src/voices.rs#L412" "open the source"
    click n_closest_pair href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-core/src/voices.rs#L432" "open the source"
    click n_distinct_voices href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-core/src/voices.rs#L458" "open the source"
    classDef entry fill:#1f2335,stroke:#7aa2f7,color:#c0caf5
    class n_applied_to,n_checked,n_describe,n_voice,n_clear_voices,n_closest_pair,n_distinct_voices entry
    classDef api fill:#1f2335,stroke:#7dcfff,color:#c0caf5
    class n_rendered_f0_hz,n_bin_hz,n_all,n_separation api

This site loads no third-party script, so it cannot run Mermaid; the diagram above is the same nodes and edges drawn by the generator instead. GitHub renders the source below directly.

ITEMS

ItemLineDocumentation
MAX_VOICES pub const129How many distinct destination voices this engine hands out.
Voice pub struct134One destination voice: the canonical values every speaker in this slot is mapped onto.
Voice::applied_to pub fn156Apply this voice to an AccentConfig.
Voice::rendered_f0_hz pub fn170The fundamental this voice will actually be rendered at, under config.
Voice::checked pub fn183Whether this voice is inside the range the engine can render usefully.
Voice::describe pub fn227A short label for an interface: "low register, narrow tract".
F0_MIN_HZ pub const253The lowest fundamental a resynthesised voice stays intelligible at.
F0_MAX_HZ pub const255The highest fundamental that still reads as a speaking register.
CENTROID_MIN_HZ pub const257The narrowest canonical vocal tract offered.
CENTROID_MAX_HZ pub const259The widest canonical vocal tract offered.
TILT_MIN_DB_OCT pub const261The steepest permitted long-term slope, in dB per octave.
TILT_MAX_DB_OCT pub const263The flattest permitted long-term slope, in dB per octave.
bin_hz pub fn270The FFT bin spacing of a configuration, in hertz.
REGISTERS_HZ const283The four registers, each a whole number of bins at the default configuration: bins 2, 3, 4 and 5 of a 1024-point frame at 48 kHz.
TRACTS const287The three vocal-tract scales, about 22 % apart.
TABLE const305The ten destination voices, in the order they are handed out.
voice pub fn326The destination voice for slot index.
all pub fn337Every destination voice, in the order they are handed out.
separation pub fn363How far apart two voices are, as the larger of their two separations.
CLEAR_SEPARATION pub const400The separation below which two voices should not be handed to two people.
clear_voices pub fn412How many voices can be handed out before two of them are too alike.
closest_pair pub fn432The closest pair among the first count voices, as a ratio.
distinct_voices pub fn458How many of the ten are still distinguishable under config.