crates/veilvoice-core/src/accent.rs
veilvoice-core · 701 lines · read the source here · or on GitHub
Accent and speaker-trait neutralisation.
What an accent is made of, and what a signal-level transform can remove
Accent is carried by two very different kinds of cue, and they have opposite answers here:
- Suprasegmental cues, meaning intonation contour and pitch range, long-term voice quality and spectral tilt, and the fixed vocal-tract scale behind a speaker's vowel space. These are properties of the signal, they are strongly speaker- and region-identifying, and this module removes them by mapping every speaker onto one canonical target.
- Segmental cues, meaning which phonemes the speaker actually produced: rhoticity, vowel mergers, dental-fricative substitution, aspiration patterns. These are not a colouration laid over the words; at this level they are the words. Removing them means deciding a different phoneme was said, which no filter can do: it requires recognising the speech and re-synthesising it (see the planned text-to-speech mode, which sidesteps this entirely by never carrying the original signal at all).
So this module makes every speaker land on the same pitch register, the same apparent vocal-tract length and the same long-term spectral tilt, which removes the accent's melody and colour and a large part of its perceived origin. It does not, and cannot, re-articulate phonemes. docs/WHITEPAPER.md must state that limit plainly rather than claim accent removal is total.
Why this also strengthens de-identification
Every step here is many-to-one: a whole population of input f0 contours, spectral tilts and vocal-tract lengths is collapsed onto a single canonical value. That destroys information rather than displacing it, so it composes with the phase discard in crate::spectral, because the two are independent one-way steps, and normalising the speaker's mean pitch and vocal-tract length removes two of the strongest biometric features there are.
Preserving intelligibility
The critical design rule is that every correction is derived from a long-term average, never from the current frame. Per-frame spectral shape is what distinguishes /i/ from /u/; normalising it frame-by-frame would erase the vowels along with the accent. Vocal-tract and tilt corrections therefore use multi-second time constants, so they track the speaker and leave the phonemes moving freely underneath.
In plain words
This is the part that works on accent, and it is careful about what it claims.
An accent is two different things at once. Some of it is in the sound: how high the voice sits, how it rises and falls, the shape of the vowels. That part can be changed here, and is.
The rest of it is in the words themselves, and in the choices somebody makes between them. No amount of altering sound touches that, because it is not in the sound. So VeilVoice says accent removal is partial, and means it.
WHAT THIS FILE CONTAINS
701 lines defining 13 functions (8 public), 3 types and 10 constants. Everything below is read out of the source, so it cannot disagree with the code.
The types it owns.
struct AccentConfigline 95 · How aggressively accent and speaker traits are normalised.struct AccentStatsline 143 · Live read-out of what the neutraliser is currently doing, for the UI.struct AccentNeutralizerline 162 · Maps any speaker onto one canonical pitch register, vocal-tract scale and long-term spectrum.
What happens when it runs. These are the ways in: public, and nothing else in this file calls them, so they are what an outside caller reaches first.
AccentNeutralizer::newline 188 · Build for a given spectrum size and sample rate.AccentNeutralizer::enabledline 217 · Whether neutralisation is active.AccentNeutralizer::statsline 222 · Live read-out for the UI.AccentNeutralizer::observeline 239 · Feed this frame's f0 estimate and update the intonation correction.AccentNeutralizer::prosody_ratioline 258 · Pitch ratio to apply to the excitation this frame.AccentNeutralizer::measure_envelopeline 269 · Measure the speaker's vocal-tract scale from the unwarped envelope and update the VTLN ratio.
reacheslog_centroidAccentNeutralizer::vtln_ratioline 288 · Formant ratio to apply to the envelope this frame.AccentNeutralizer::shapeline 309 · Rotate the already-warped envelope toward the canonical spectral tilt, then fold the result back into the running average.
reachesdb_to_gain,gain_to_db,recompute_shape
WHAT CALLS WHAT
The functions this file defines, and the calls between them. An edge means the callee's name appears, called, inside the caller's body. This is a syntactic reading, not a type-resolved one.
The same graph as Mermaid source
%%{init: {"theme":"base","themeVariables":{"background":"#1a1b26","primaryColor":"#1f2335","primaryTextColor":"#c0caf5","primaryBorderColor":"#7aa2f7","secondaryColor":"#16161e","tertiaryColor":"#16161e","lineColor":"#737aa2","textColor":"#c0caf5","mainBkg":"#1f2335","nodeBorder":"#7aa2f7","clusterBkg":"#16161e","clusterBorder":"#2f3549","fontFamily":"ui-monospace, SFMono-Regular, Consolas, monospace","fontSize":"14px"}}}%%
flowchart TD
n_default["AccentConfig::default<br/>line 125"]
n_new(["AccentNeutralizer::new<br/>line 188"])
n_enabled(["AccentNeutralizer::enabled<br/>line 217"])
n_stats(["AccentNeutralizer::stats<br/>line 222"])
n_observe(["AccentNeutralizer::observe<br/>line 239"])
n_prosody_ratio(["AccentNeutralizer::<br/>prosody_ratio<br/>line 258"])
n_measure_envelope(["AccentNeutralizer::<br/>measure_envelope<br/>line 269"])
n_vtln_ratio(["AccentNeutralizer::vtln_ratio<br/>line 288"])
n_shape(["AccentNeutralizer::shape<br/>line 309"])
n_recompute_shape["AccentNeutralizer::<br/>recompute_shape<br/>line 352"]
n_log_centroid["log_centroid<br/>line 404"]
n_gain_to_db["gain_to_db<br/>line 425"]
n_db_to_gain["db_to_gain<br/>line 431"]
n_measure_envelope --> n_log_centroid
n_shape --> n_db_to_gain
n_shape --> n_gain_to_db
n_shape --> n_recompute_shape
click n_default href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-core/src/accent.rs#L125" "open the source"
click n_new href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-core/src/accent.rs#L188" "open the source"
click n_enabled href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-core/src/accent.rs#L217" "open the source"
click n_stats href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-core/src/accent.rs#L222" "open the source"
click n_observe href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-core/src/accent.rs#L239" "open the source"
click n_prosody_ratio href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-core/src/accent.rs#L258" "open the source"
click n_measure_envelope href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-core/src/accent.rs#L269" "open the source"
click n_vtln_ratio href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-core/src/accent.rs#L288" "open the source"
click n_shape href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-core/src/accent.rs#L309" "open the source"
click n_recompute_shape href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-core/src/accent.rs#L352" "open the source"
click n_log_centroid href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-core/src/accent.rs#L404" "open the source"
click n_gain_to_db href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-core/src/accent.rs#L425" "open the source"
click n_db_to_gain href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-core/src/accent.rs#L431" "open the source"
classDef entry fill:#1f2335,stroke:#7aa2f7,color:#c0caf5
class n_new,n_enabled,n_stats,n_observe,n_prosody_ratio,n_measure_envelope,n_vtln_ratio,n_shape entry
classDef helper fill:#1f2335,stroke:#bb9af7,color:#c0caf5
class n_default,n_recompute_shape,n_log_centroid,n_gain_to_db,n_db_to_gain helper
This site loads no third-party script, so it cannot run Mermaid; the diagram above is the same nodes and edges drawn by the generator instead. GitHub renders the source below directly.
ITEMS
| Item | Line | Documentation |
|---|---|---|
CENTROID_LO_HZ const | 62 | Lower edge of the band used to measure vocal-tract scale, in hertz. |
CENTROID_HI_HZ const | 64 | Upper edge of the band used to measure vocal-tract scale, in hertz. |
LTAS_LO_HZ const | 66 | Lower edge of the band whose long-term tilt is normalised. |
LTAS_HI_HZ const | 68 | Upper edge of the band whose long-term tilt is normalised. |
TILT_REF_HZ const | 70 | Reference frequency of the canonical spectral-tilt line, in hertz. |
MAX_SHAPE_DB const | 72 | Maximum long-term shaping applied to any bin, in decibels. |
TAU_PROSODY_S const | 74 | Time constant for the intonation correction, in seconds. |
TAU_VTLN_S const | 77 | Time constant for the vocal-tract estimate, in seconds. |
TAU_LTAS_S const | 79 | Time constant for the long-term average spectrum, in seconds. |
WARMUP_S pub const | 88 | Seconds of voiced audio over which corrections fade in from nothing. |
AccentConfig pub struct | 95 | How aggressively accent and speaker traits are normalised. |
AccentConfig::default fn | 125 | |
AccentStats pub struct | 143 | Live read-out of what the neutraliser is currently doing, for the UI. |
AccentNeutralizer pub struct | 162 | Maps any speaker onto one canonical pitch register, vocal-tract scale and long-term spectrum. |
AccentNeutralizer::new pub fn | 188 | Build for a given spectrum size and sample rate. |
AccentNeutralizer::enabled pub fn | 217 | Whether neutralisation is active. |
AccentNeutralizer::stats pub fn | 222 | Live read-out for the UI. |
AccentNeutralizer::observe pub fn | 239 | Feed this frame's f0 estimate and update the intonation correction. |
AccentNeutralizer::prosody_ratio pub fn | 258 | Pitch ratio to apply to the excitation this frame. |
AccentNeutralizer::measure_envelope pub fn | 269 | Measure the speaker's vocal-tract scale from the unwarped envelope and update the VTLN ratio. |
AccentNeutralizer::vtln_ratio pub fn | 288 | Formant ratio to apply to the envelope this frame. |
AccentNeutralizer::shape pub fn | 309 | Rotate the already-warped envelope toward the canonical spectral tilt, then fold the result back into the running average. |
AccentNeutralizer::recompute_shape fn | 352 | Rebuild the correction curve from the current long-term average. |
log_centroid fn | 404 | Energy-weighted geometric-mean frequency of env over lo, hi bins. |
gain_to_db fn | 425 | A linear gain as decibels. |
db_to_gain fn | 431 | Decibels back to a linear gain. |