crates/veilvoice-conversation/src/plan.rs
veilvoice-conversation · 1201 lines · read the source here · or on GitHub
Who is in the recording, and who is speaking when.
VeilVoice does not work out who is talking, and will not guess
Deciding which person is speaking at each moment is speaker diarisation, and doing it from the audio alone needs a trained model. This project ships no model, talks to no server, and is not about to start doing either, so the turns come from the user, and there are exactly two honest ways to get them:
- One microphone each. If the recording has a channel per person, the split is already there and is exact.
Conversation::from_channelsbuilds the plan from that. - A list of turns. Times and speakers, in a text file, written by whoever was there or produced by whatever tool they already use for transcripts.
What would be worse than either is guessing. A wrong guess maps two people onto one voice, which is a privacy improvement and a usability disaster, or splits one person across two voices, which invites a listener to believe there was somebody in the room who was not. Neither failure would be visible in the output, and both would be blamed on the recording rather than on the tool.
Format
Text, one record per line, for the same reason everything else here is text: a file describing who said what is worth more if it can be read, checked and edited without this program.
VEILCONV1
title Two people, one microphone
speaker 0 Alex
speaker 1 Sam portrait.png
turn 0.000 4.200 0 Hello -- how did it go?
turn 4.100 9.050 1
Times are seconds with a decimal point. The text on a turn is optional: with it, subtitles carry the words; without it they carry the speaker's name and nothing else, which is still enough to follow a conversation whose voices have all been replaced.
Overlapping turns are allowed, because people talk over each other, and crate::render mixes them rather than picking a winner.
In plain words
A list of who is in a recording and when each of them speaks.
VeilVoice does not work this out for itself. Deciding who is talking at any moment is a hard problem that needs a trained model, and this project does not ship one, so it asks instead. You either write the times down, or you record each person on their own microphone.
That is less convenient and it is honest. A program that guessed would sometimes put one person's words in another person's voice, and you would not find out by listening, because the result would sound perfectly fine.
WHAT THIS FILE CONTAINS
1201 lines defining 25 functions (22 public), 3 types and 1 constant. Everything below is read out of the source, so it cannot disagree with the code.
The types it owns.
struct Speakerline 71 · One person in the recording.struct Turnline 112 · A span of the recording belonging to one speaker.struct Conversationline 136 · The whole plan: who is in the recording, and when each of them speaks.
What happens when it runs. These are the ways in: public, and nothing else in this file calls them, so they are what an outside caller reaches first.
Turn::durationline 129 · How long this turn lasts, in seconds.Conversation::add_speakerline 172 · Add a speaker, and return the index they were given.Conversation::add_turnline 202 · Add a turn.Conversation::rename_speakersline 268 · Rename everybody, in slot order, keeping every turn where it is.Conversation::speakersline 298 · The speakers, in the order they were added.Conversation::turnsline 303 · The turns, in time order.Conversation::lenline 308 · How many speakers there are.Conversation::is_emptyline 313 · Whether there is nobody in the plan.Conversation::colour_ofline 344 · The colour to draw a speaker in, chosen or from the palette.Conversation::voiceline 352 · The destination voice for a speaker.Conversation::modeline 363 · Whether every speaker gets their own voice, or one between them.Conversation::set_modeline 372 · Render every speaker as the same voice, or as their own.Conversation::durationline 383 · When the last turn ends, in seconds.Conversation::overlapsline 392 · Turns where two people are speaking at once.Conversation::self_overlapsline 413 · Spans where a speaker's turns overlap their own other turns.Conversation::from_channelsline 435 · A plan for a recording with one microphone per person.
reachesnamed,newConversation::saveline 652 · Write the plan to path.
reachesto_textConversation::loadline 668 · Read a plan written by Conversation::save.
reachesparse,new,split_word
WHAT CALLS WHAT
The functions this file defines, and the calls between them. An edge means the callee's name appears, called, inside the caller's body. This is a syntactic reading, not a type-resolved one. 22 of 24 functions are drawn; the diagram is bounded at 22 so it stays readable.
The same graph as Mermaid source
%%{init: {"theme":"base","themeVariables":{"background":"#1a1b26","primaryColor":"#1f2335","primaryTextColor":"#c0caf5","primaryBorderColor":"#7aa2f7","secondaryColor":"#16161e","tertiaryColor":"#16161e","lineColor":"#737aa2","textColor":"#c0caf5","mainBkg":"#1f2335","nodeBorder":"#7aa2f7","clusterBkg":"#16161e","clusterBorder":"#2f3549","fontFamily":"ui-monospace, SFMono-Regular, Consolas, monospace","fontSize":"14px"}}}%%
flowchart TD
n_named["Speaker::named<br/>line 101"]
n_duration(["Turn::duration<br/>line 129"])
n_split_word["split_word<br/>line 154"]
n_new["Conversation::new<br/>line 162"]
n_add_speaker(["Conversation::add_speaker<br/>line 172"])
n_add_turn(["Conversation::add_turn<br/>line 202"])
n_rename_speakers(["Conversation::rename_speakers<br/>line 268"])
n_speakers(["Conversation::speakers<br/>line 298"])
n_turns(["Conversation::turns<br/>line 303"])
n_len(["Conversation::len<br/>line 308"])
n_is_empty(["Conversation::is_empty<br/>line 313"])
n_colour_of(["Conversation::colour_of<br/>line 344"])
n_voice(["Conversation::voice<br/>line 352"])
n_mode(["Conversation::mode<br/>line 363"])
n_set_mode(["Conversation::set_mode<br/>line 372"])
n_duration(["Conversation::duration<br/>line 383"])
n_overlaps(["Conversation::overlaps<br/>line 392"])
n_from_channels(["Conversation::from_channels<br/>line 435"])
n_to_text["Conversation::to_text<br/>line 460"]
n_parse["Conversation::parse<br/>line 504"]
n_save(["Conversation::save<br/>line 652"])
n_load(["Conversation::load<br/>line 668"])
n_from_channels --> n_named
n_from_channels --> n_new
n_load --> n_parse
n_parse --> n_new
n_parse --> n_split_word
n_save --> n_to_text
click n_named href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-conversation/src/plan.rs#L101" "open the source"
click n_duration href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-conversation/src/plan.rs#L129" "open the source"
click n_split_word href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-conversation/src/plan.rs#L154" "open the source"
click n_new href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-conversation/src/plan.rs#L162" "open the source"
click n_add_speaker href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-conversation/src/plan.rs#L172" "open the source"
click n_add_turn href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-conversation/src/plan.rs#L202" "open the source"
click n_rename_speakers href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-conversation/src/plan.rs#L268" "open the source"
click n_speakers href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-conversation/src/plan.rs#L298" "open the source"
click n_turns href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-conversation/src/plan.rs#L303" "open the source"
click n_len href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-conversation/src/plan.rs#L308" "open the source"
click n_is_empty href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-conversation/src/plan.rs#L313" "open the source"
click n_colour_of href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-conversation/src/plan.rs#L344" "open the source"
click n_voice href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-conversation/src/plan.rs#L352" "open the source"
click n_mode href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-conversation/src/plan.rs#L363" "open the source"
click n_set_mode href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-conversation/src/plan.rs#L372" "open the source"
click n_duration href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-conversation/src/plan.rs#L383" "open the source"
click n_overlaps href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-conversation/src/plan.rs#L392" "open the source"
click n_from_channels href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-conversation/src/plan.rs#L435" "open the source"
click n_to_text href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-conversation/src/plan.rs#L460" "open the source"
click n_parse href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-conversation/src/plan.rs#L504" "open the source"
click n_save href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-conversation/src/plan.rs#L652" "open the source"
click n_load href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-conversation/src/plan.rs#L668" "open the source"
classDef entry fill:#1f2335,stroke:#7aa2f7,color:#c0caf5
class n_duration,n_add_speaker,n_add_turn,n_rename_speakers,n_speakers,n_turns,n_len,n_is_empty,n_colour_of,n_voice,n_mode,n_set_mode,n_duration,n_overlaps,n_from_channels,n_save,n_load entry
classDef api fill:#1f2335,stroke:#7dcfff,color:#c0caf5
class n_named,n_new,n_to_text,n_parse api
classDef helper fill:#1f2335,stroke:#bb9af7,color:#c0caf5
class n_split_word helper
This site loads no third-party script, so it cannot run Mermaid; the diagram above is the same nodes and edges drawn by the generator instead. GitHub renders the source below directly.
ITEMS
| Item | Line | Documentation |
|---|---|---|
MAGIC const | 67 | Magic first line. |
Speaker pub struct | 71 | One person in the recording. |
Speaker::named pub fn | 101 | A speaker with a name and no picture. |
Turn pub struct | 112 | A span of the recording belonging to one speaker. |
Turn::duration pub fn | 129 | How long this turn lasts, in seconds. |
Conversation pub struct | 136 | The whole plan: who is in the recording, and when each of them speaks. |
split_word fn | 154 | The first whitespace-separated word, and everything after it. |
Conversation::new pub fn | 162 | An empty plan. |
Conversation::add_speaker pub fn | 172 | Add a speaker, and return the index they were given. |
Conversation::add_turn pub fn | 202 | Add a turn. |
Conversation::rename_speakers pub fn | 268 | Rename everybody, in slot order, keeping every turn where it is. |
Conversation::speakers pub fn | 298 | The speakers, in the order they were added. |
Conversation::turns pub fn | 303 | The turns, in time order. |
Conversation::len pub fn | 308 | How many speakers there are. |
Conversation::is_empty pub fn | 313 | Whether there is nobody in the plan. |
Conversation::speakers_mut pub(crate) fn | 324 | The speakers, for an edit that has already checked what it is doing. |
Conversation::turns_mut pub(crate) fn | 335 | The spans, for an edit that has already checked what it is doing. |
Conversation::colour_of pub fn | 344 | The colour to draw a speaker in, chosen or from the palette. |
Conversation::voice pub fn | 352 | The destination voice for a speaker. |
Conversation::mode pub fn | 363 | Whether every speaker gets their own voice, or one between them. |
Conversation::set_mode pub fn | 372 | Render every speaker as the same voice, or as their own. |
Conversation::duration pub fn | 383 | When the last turn ends, in seconds. |
Conversation::overlaps pub fn | 392 | Turns where two people are speaking at once. |
Conversation::self_overlaps pub fn | 413 | Spans where a speaker's turns overlap their own other turns. |
Conversation::from_channels pub fn | 435 | A plan for a recording with one microphone per person. |
Conversation::to_text pub fn | 460 | Serialise to the text format described at the top of this module. |
Conversation::parse pub fn | 504 | Parse the text format. |
Conversation::save pub fn | 652 | Write the plan to path. |
Conversation::load pub fn | 668 | Read a plan written by Conversation::save. |
guide_tests mod | 1104 |