plan.rs

crates/veilvoice-conversation/src/plan.rs

veilvoice-conversation · 1201 lines · read the source here · or on GitHub

Who is in the recording, and who is speaking when.

VeilVoice does not work out who is talking, and will not guess

Deciding which person is speaking at each moment is speaker diarisation, and doing it from the audio alone needs a trained model. This project ships no model, talks to no server, and is not about to start doing either, so the turns come from the user, and there are exactly two honest ways to get them:

  • One microphone each. If the recording has a channel per person, the split is already there and is exact. Conversation::from_channels builds the plan from that.
  • A list of turns. Times and speakers, in a text file, written by whoever was there or produced by whatever tool they already use for transcripts.

What would be worse than either is guessing. A wrong guess maps two people onto one voice, which is a privacy improvement and a usability disaster, or splits one person across two voices, which invites a listener to believe there was somebody in the room who was not. Neither failure would be visible in the output, and both would be blamed on the recording rather than on the tool.

Format

Text, one record per line, for the same reason everything else here is text: a file describing who said what is worth more if it can be read, checked and edited without this program.

VEILCONV1
title  Two people, one microphone
speaker  0  Alex
speaker  1  Sam  portrait.png
turn  0.000  4.200  0  Hello -- how did it go?
turn  4.100  9.050  1

Times are seconds with a decimal point. The text on a turn is optional: with it, subtitles carry the words; without it they carry the speaker's name and nothing else, which is still enough to follow a conversation whose voices have all been replaced.

Overlapping turns are allowed, because people talk over each other, and crate::render mixes them rather than picking a winner.

In plain words

A list of who is in a recording and when each of them speaks.

VeilVoice does not work this out for itself. Deciding who is talking at any moment is a hard problem that needs a trained model, and this project does not ship one, so it asks instead. You either write the times down, or you record each person on their own microphone.

That is less convenient and it is honest. A program that guessed would sometimes put one person's words in another person's voice, and you would not find out by listening, because the result would sound perfectly fine.

WHAT THIS FILE CONTAINS

1201 lines defining 25 functions (22 public), 3 types and 1 constant. Everything below is read out of the source, so it cannot disagree with the code.

The types it owns.

  • struct Speaker line 71 · One person in the recording.
  • struct Turn line 112 · A span of the recording belonging to one speaker.
  • struct Conversation line 136 · The whole plan: who is in the recording, and when each of them speaks.

What happens when it runs. These are the ways in: public, and nothing else in this file calls them, so they are what an outside caller reaches first.

  • Turn::duration line 129 · How long this turn lasts, in seconds.
  • Conversation::add_speaker line 172 · Add a speaker, and return the index they were given.
  • Conversation::add_turn line 202 · Add a turn.
  • Conversation::rename_speakers line 268 · Rename everybody, in slot order, keeping every turn where it is.
  • Conversation::speakers line 298 · The speakers, in the order they were added.
  • Conversation::turns line 303 · The turns, in time order.
  • Conversation::len line 308 · How many speakers there are.
  • Conversation::is_empty line 313 · Whether there is nobody in the plan.
  • Conversation::colour_of line 344 · The colour to draw a speaker in, chosen or from the palette.
  • Conversation::voice line 352 · The destination voice for a speaker.
  • Conversation::mode line 363 · Whether every speaker gets their own voice, or one between them.
  • Conversation::set_mode line 372 · Render every speaker as the same voice, or as their own.
  • Conversation::duration line 383 · When the last turn ends, in seconds.
  • Conversation::overlaps line 392 · Turns where two people are speaking at once.
  • Conversation::self_overlaps line 413 · Spans where a speaker's turns overlap their own other turns.
  • Conversation::from_channels line 435 · A plan for a recording with one microphone per person.
    reaches named, new
  • Conversation::save line 652 · Write the plan to path.
    reaches to_text
  • Conversation::load line 668 · Read a plan written by Conversation::save.
    reaches parse, new, split_word

WHAT CALLS WHAT

The functions this file defines, and the calls between them. An edge means the callee's name appears, called, inside the caller's body. This is a syntactic reading, not a type-resolved one. 22 of 24 functions are drawn; the diagram is bounded at 22 so it stays readable.

The same graph as Mermaid source
%%{init: {"theme":"base","themeVariables":{"background":"#1a1b26","primaryColor":"#1f2335","primaryTextColor":"#c0caf5","primaryBorderColor":"#7aa2f7","secondaryColor":"#16161e","tertiaryColor":"#16161e","lineColor":"#737aa2","textColor":"#c0caf5","mainBkg":"#1f2335","nodeBorder":"#7aa2f7","clusterBkg":"#16161e","clusterBorder":"#2f3549","fontFamily":"ui-monospace, SFMono-Regular, Consolas, monospace","fontSize":"14px"}}}%%
flowchart TD
    n_named["Speaker::named<br/>line 101"]
    n_duration(["Turn::duration<br/>line 129"])
    n_split_word["split_word<br/>line 154"]
    n_new["Conversation::new<br/>line 162"]
    n_add_speaker(["Conversation::add_speaker<br/>line 172"])
    n_add_turn(["Conversation::add_turn<br/>line 202"])
    n_rename_speakers(["Conversation::rename_speakers<br/>line 268"])
    n_speakers(["Conversation::speakers<br/>line 298"])
    n_turns(["Conversation::turns<br/>line 303"])
    n_len(["Conversation::len<br/>line 308"])
    n_is_empty(["Conversation::is_empty<br/>line 313"])
    n_colour_of(["Conversation::colour_of<br/>line 344"])
    n_voice(["Conversation::voice<br/>line 352"])
    n_mode(["Conversation::mode<br/>line 363"])
    n_set_mode(["Conversation::set_mode<br/>line 372"])
    n_duration(["Conversation::duration<br/>line 383"])
    n_overlaps(["Conversation::overlaps<br/>line 392"])
    n_from_channels(["Conversation::from_channels<br/>line 435"])
    n_to_text["Conversation::to_text<br/>line 460"]
    n_parse["Conversation::parse<br/>line 504"]
    n_save(["Conversation::save<br/>line 652"])
    n_load(["Conversation::load<br/>line 668"])
    n_from_channels --> n_named
    n_from_channels --> n_new
    n_load --> n_parse
    n_parse --> n_new
    n_parse --> n_split_word
    n_save --> n_to_text
    click n_named href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-conversation/src/plan.rs#L101" "open the source"
    click n_duration href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-conversation/src/plan.rs#L129" "open the source"
    click n_split_word href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-conversation/src/plan.rs#L154" "open the source"
    click n_new href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-conversation/src/plan.rs#L162" "open the source"
    click n_add_speaker href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-conversation/src/plan.rs#L172" "open the source"
    click n_add_turn href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-conversation/src/plan.rs#L202" "open the source"
    click n_rename_speakers href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-conversation/src/plan.rs#L268" "open the source"
    click n_speakers href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-conversation/src/plan.rs#L298" "open the source"
    click n_turns href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-conversation/src/plan.rs#L303" "open the source"
    click n_len href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-conversation/src/plan.rs#L308" "open the source"
    click n_is_empty href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-conversation/src/plan.rs#L313" "open the source"
    click n_colour_of href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-conversation/src/plan.rs#L344" "open the source"
    click n_voice href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-conversation/src/plan.rs#L352" "open the source"
    click n_mode href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-conversation/src/plan.rs#L363" "open the source"
    click n_set_mode href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-conversation/src/plan.rs#L372" "open the source"
    click n_duration href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-conversation/src/plan.rs#L383" "open the source"
    click n_overlaps href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-conversation/src/plan.rs#L392" "open the source"
    click n_from_channels href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-conversation/src/plan.rs#L435" "open the source"
    click n_to_text href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-conversation/src/plan.rs#L460" "open the source"
    click n_parse href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-conversation/src/plan.rs#L504" "open the source"
    click n_save href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-conversation/src/plan.rs#L652" "open the source"
    click n_load href "https://github.com/tilas01/veilvoice/blob/main/crates/veilvoice-conversation/src/plan.rs#L668" "open the source"
    classDef entry fill:#1f2335,stroke:#7aa2f7,color:#c0caf5
    class n_duration,n_add_speaker,n_add_turn,n_rename_speakers,n_speakers,n_turns,n_len,n_is_empty,n_colour_of,n_voice,n_mode,n_set_mode,n_duration,n_overlaps,n_from_channels,n_save,n_load entry
    classDef api fill:#1f2335,stroke:#7dcfff,color:#c0caf5
    class n_named,n_new,n_to_text,n_parse api
    classDef helper fill:#1f2335,stroke:#bb9af7,color:#c0caf5
    class n_split_word helper

This site loads no third-party script, so it cannot run Mermaid; the diagram above is the same nodes and edges drawn by the generator instead. GitHub renders the source below directly.

ITEMS

ItemLineDocumentation
MAGIC const67Magic first line.
Speaker pub struct71One person in the recording.
Speaker::named pub fn101A speaker with a name and no picture.
Turn pub struct112A span of the recording belonging to one speaker.
Turn::duration pub fn129How long this turn lasts, in seconds.
Conversation pub struct136The whole plan: who is in the recording, and when each of them speaks.
split_word fn154The first whitespace-separated word, and everything after it.
Conversation::new pub fn162An empty plan.
Conversation::add_speaker pub fn172Add a speaker, and return the index they were given.
Conversation::add_turn pub fn202Add a turn.
Conversation::rename_speakers pub fn268Rename everybody, in slot order, keeping every turn where it is.
Conversation::speakers pub fn298The speakers, in the order they were added.
Conversation::turns pub fn303The turns, in time order.
Conversation::len pub fn308How many speakers there are.
Conversation::is_empty pub fn313Whether there is nobody in the plan.
Conversation::speakers_mut pub(crate) fn324The speakers, for an edit that has already checked what it is doing.
Conversation::turns_mut pub(crate) fn335The spans, for an edit that has already checked what it is doing.
Conversation::colour_of pub fn344The colour to draw a speaker in, chosen or from the palette.
Conversation::voice pub fn352The destination voice for a speaker.
Conversation::mode pub fn363Whether every speaker gets their own voice, or one between them.
Conversation::set_mode pub fn372Render every speaker as the same voice, or as their own.
Conversation::duration pub fn383When the last turn ends, in seconds.
Conversation::overlaps pub fn392Turns where two people are speaking at once.
Conversation::self_overlaps pub fn413Spans where a speaker's turns overlap their own other turns.
Conversation::from_channels pub fn435A plan for a recording with one microphone per person.
Conversation::to_text pub fn460Serialise to the text format described at the top of this module.
Conversation::parse pub fn504Parse the text format.
Conversation::save pub fn652Write the plan to path.
Conversation::load pub fn668Read a plan written by Conversation::save.
guide_tests mod1104