Learning how to Label a Church Service
Can a model that never hears the audio tell a sermon from a song?
At transportin.vesper.audio we take a full recording of a church service and draw a colored timeline under the waveform: this stretch is worship, this is the sermon, this is the closing song. You can drop in your own file and get the same timeline back a few minutes later. The question behind the whole feature is simple to state: given 92 minutes of a broadcast mix, where does each part of the service begin and end?
There are eight kinds of moment we care about, and the output is just a list of
{start, end, label} entries covering the whole recording:
The obvious way to build this is to train an audio model: collect thousands of labeled services, learn what worship sounds like versus a sermon, ship a classifier. We did not do that. Transportin's classifier has no trained model of its own at all, and this post explains what it does instead: it turns the audio into a transcript, asks a judgment API called Jev one multiple-choice question per 30 seconds, and then runs a 60-year-old optimal-path algorithm over the answers. Each piece is simple. The interesting part is why the pieces are shaped the way they are.
A classifier that never hears the audio
The pipeline has two halves. The first half is perception: an ASR model
(NVIDIA's parakeet-tdt-0.6b-v3, running on a GPU on Modal) chews through the audio
in 10-minute chunks and produces a transcript with a timestamp on every word.
For our 92-minute sample that's about 19,000 timed words. The second half is judgment: deciding,
for each moment, which of the eight categories that moment belongs to. Pictorially:
Jev (the model behind TypeSafe AI's API) is not a chat model and you don't prompt it with prose. Every call has the same shape: you send a state (any JSON object describing a situation) and a question about that state. For a choice question you also send the allowed answers, called criteria, each with a plain-English definition. Jev must answer with one of your criteria, and it returns a probability for every criterion, not just its pick. Here is the actual request body transportin sends, abridged:
{
"state": { ...the 30-second window, described below... },
"model": "jev-latest",
"questions": {
"segment": {
"type": "choice",
"instructions": "This is one time-window from the broadcast audio mix
of a church service. Which segment type is this window part of?
Weigh the transcript text, how much speech there is (sparse or no
words usually means music, a video bed, or an empty room), and
where the window falls within the whole service.",
"criteria": {
"worship": "Live congregational worship music: the band playing and
singing song lyrics, repeated sung phrases of praise; no extended
talking.",
"sermon": "The preaching or teaching block: one speaker in a
sustained monologue, scripture reading and exposition,
illustrations, application.",
...six more definitions...
}
}
}
}
And a real response, for the window starting at 29:30 in the sample1:
{ "answers": { "segment": {
"type": "choice",
"choice": "sermon",
"confidence": 0.72,
"probabilities": { "sermon": 0.75, "mc": 0.14, "worship_hybrid": 0.09,
"worship": 0.01, "video": 0.01, "pre_roll": 0.0,
"closing_song": 0.0, "post_roll": 0.0 }
} } }
Notice what this design buys us. There is no prompt template to babysit and no free-text output
to parse: the "model" for our problem is literally the eight criteria definitions, about 90
words of English. Want a ninth category? Add one dictionary entry. Want to sharpen the boundary
between worship and worship_hybrid? Edit two sentences. The entire
classifier is legible in a way a fine-tuned audio model never is.
What does 30 seconds of church look like as JSON?
The transcript is cut into 30-second windows (184 of them for the sample file), and each window
becomes one state object. Here is the real state for that 29:30 window, lightly
trimmed:
{
"context": "One window of broadcast audio from a full church service.",
"window": "00:29:30 to 00:30:00",
"service_duration": "01:31:57",
"position_in_service_pct": 32.1,
"words_per_minute": 214.0,
"transcript": "This is great. Oh my lord. But uh but yeah, no, it was
like dad had me prepared. He always was like, Yeah, you gotta know it
better than anyone... [capped at 1,500 characters]",
"previous_window_tail": "...last 300 characters of the window before...",
"next_window_head": "...first 300 characters of the window after..."
}
Every field is there for a reason, and the reasons matter more than the fields:
- The transcript is the main evidence. Song lyrics, announcements, and preaching read very differently on the page, and Jev is good at telling them apart. Empty windows send the literal string "(no transcribed speech in this window)" rather than nothing, so silence is evidence too.
- Words-per-minute is the music detector. This is the crucial trick, because Jev is deaf. It never gets a spectrogram, an energy curve, or a single audio feature. But an ASR transcript of a band playing runs at a fraction of the word rate of a preacher mid-sermon (this window's 214 wpm is unambiguous talking). Sparse words mean band, video bed, or an empty room, and the instructions tell Jev to reason exactly that way.
- Position in the service breaks text ties. A song at 5% of the way
through and a song at 95% can have identical transcripts, but one is
worshipand the other isclosing_song. Same forpre_rollversuspost_roll: both are "music, nobody talking to the room," distinguished almost entirely by the clock. - The neighbor snippets are cheap context. 300 characters from each side is enough for Jev to see that a quiet window sits inside a sermon rather than between songs, without paying to send three full windows.
So each Jev call is a self-contained question: here is half a minute of a service, in text and numbers; which of these eight things is it? Asked 184 times, in parallel, eight at a time.
The raw answers are jumpy
If we simply trusted each window's top pick, we'd have a problem. Here is the entire sample run, one thin bar per window, raw picks on top:
The raw row would give us 56 segments in 92 minutes. Nobody's service has 56 segments. The flicker isn't Jev being bad at its job; it's Jev answering exactly the question we asked. Each call sees one window in isolation, and some individual half-minutes genuinely are ambiguous: the preacher tells a joke and the window reads like an emcee; there's a burst of crosstalk and it reads like a video. Window-level ambiguity is unavoidable. Service-level flicker is not, because we know something the per-window question can't express: services are contiguous. A sermon is twenty consecutive minutes, not a strobe light of sermon-mc-sermon-mc.
Making a label change expensive
The fix is to stop treating the 184 answers as 184 decisions and make one global decision instead. First, turn each probability into a cost:
Confident answers become cheap labels, doubtful answers become expensive ones. A probability of 0.75 costs 0.29; a probability of 0.11 costs 2.2.2
Then define the best labeling recursively: the cheapest way to label windows 0 through i ending in label ℓ is the cost of putting ℓ on window i, plus the cheapest way to have labeled everything before it, plus a flat penalty if the previous window's label was different:
The switch penalty of 2.5 is the whole personality of the smoother. Solving this recurrence for all windows and taking the cheapest final state is the Viterbi algorithm: dynamic programming over an optimal-path problem, the same trick that powers everything from GPS routing to speech recognition.
What does a penalty of 2.5 mean in human terms? Costs live in log space, so adding 2.5 is the same as dividing probability by e2.5 ≈ 12. A label change has to buy back a 12× probability advantage somewhere to be worth making. And a one-window island is even harder to justify: it needs a switch on the way in and a switch on the way out, a toll of 5.0, or roughly a 150× evidence bar.
Watch it work on a real case. At 1:30 in the sample there's a burst of produced-sounding audio,
and Jev's raw pick for that window is video at 0.45, with mc second at
0.23:
video island would gain
log(0.45/0.23) ≈ 0.7 of evidence but pay 5.0 in switch tolls. The smoother
relabels it mc, and the surrounding announcement block stays whole.
The evidence for video over mc in that window is about 2×.
The bar for a one-window island is about 150×. Overruled. But this is not a
"majority wins" hack: if six consecutive windows had said video at those odds, the
island would pay its two tolls once and collect the evidence six times, and it would survive.
The penalty doesn't forbid short segments; it demands they earn their keep. Across the sample
run, smoothing overrode Jev's raw pick on 36 of 184 windows, and the timeline
went from 56 segments to 8.
From windows to segments
The last step is bookkeeping. Consecutive windows that share a smoothed label merge into one segment, and each segment's confidence is the average of Jev's probability for the winning label across its windows. Here is the actual shipped output for the sample file:
| start | end | label | confidence |
|---|---|---|---|
| 00:00 | 01:30 | sermon | 0.75 |
| 01:30 | 13:30 | mc | 0.60 |
| 13:30 | 18:00 | sermon | 0.69 |
| 18:00 | 36:30 | mc | 0.46 |
| 36:30 | 1:13:00 | sermon | 0.65 |
| 1:13:00 | 1:23:00 | mc | 0.42 |
| 1:23:00 | 1:30:00 | sermon | 0.54 |
| 1:30:00 | 1:31:57 | mc | 0.39 |
Note the 0.46s and 0.42s. Those numbers are honest: they say Jev was genuinely torn across those stretches, and the smoother made the call the individual windows couldn't. A confidence near 0.9 means the windows agreed on their own; a confidence near 0.4 means the timeline is leaning on contiguity. Both are useful things for a reader of the timeline to know.
And why only two colors? Because the public sample is, frankly, the wrong audio: it's a
podcast conversation between musicians, not a church service. Of 184 raw picks, 96 said
sermon (sustained monologue), 77 said mc (conversational back and
forth), and only 11 said anything else. That is the system working correctly on material
where six of its eight categories can't exist. On a real broadcast mix all eight fire; a real
service can't be published on the public site without the church's say-so.
What it costs to run
The unit of spend is one Jev call per 30-second window: about 120 calls per hour of audio, fired eight at a time from a thread pool, with exponential backoff on rate limits. The pipeline runs identically in two places (a local CLI for the baked-in sample, and a Modal function for public uploads) and it is deliberately the only code that talks to Jev. The web server never holds the API key; it just kicks off a job and polls a status file.
One design choice pays for itself daily: every raw Jev answer is cached to disk, keyed by window index3, and the smoothing runs after the cache. So retuning the switch penalty, the knob you actually want to play with, costs zero API calls. You can re-run the whole smoother in under a second, compare 8 segments against 12, and never touch the network. The expensive, slow, rate-limited part of the pipeline runs once; the judgment-shaping part iterates freely.
Final thoughts
The pattern here generalizes well past church audio. Instead of training a bespoke classifier, we (1) reduced the raw signal to text plus a few cheap numbers, (2) asked a general judgment model a tightly-scoped multiple-choice question many times, keeping full distributions, and (3) recovered the structure we knew about, contiguity, with classical dynamic programming rather than more machine learning. The eight criteria definitions do the work a training set would, at one-millionth the effort, and every stage of the pipeline can be inspected by reading English or a 20-line algorithm.
One honest limitation to end on. The site now has a live "streaming segmentation" page, and it is a UI demo only: the segments there are simulated, and Jev is not connected to it. The per-window Jev call would stream just fine. The smoother would not: Viterbi as written is non-causal, meaning it reads the whole service before labeling the first window (that's exactly where its power comes from). A live version has to trade some of that power for latency, deciding each window after seeing only a few windows beyond it, and tuning that lag against the 12× evidence bar is the genuinely interesting open problem between the demo and the real thing.
1 Answers come wrapped in an answers
envelope keyed by your question's name, not at the top level of the response. This cost the
first implementation a KeyError. ↩
2 Probabilities are floored at 10−6 before the log, so a hard 0.0 from Jev becomes a very expensive label instead of an infinitely forbidden one. Without the floor, one overconfident zero could make the true label unreachable for the whole service. ↩
3 The cache key includes the window size
(w30.0:59), so changing the windowing re-classifies everything (correctly,
since every window's text changes) while changing only the smoothing reuses every
cached answer. ↩