A single timed transcript cue — the unit produced by ASR and by MJ's realtime session capture (which records speaker + timings per turn).
Cue end, milliseconds from the beginning of the asset.
Optional
Optional speaker label/id — a speaker change is a strong boundary signal.
Cue start, milliseconds from the beginning of the asset.
Spoken text for this cue.
A single timed transcript cue — the unit produced by ASR and by MJ's realtime session capture (which records speaker + timings per turn).