Hundreds of hours of rushes, and no reliable way to know what is in them.
A large-cast unscripted series comes off the floor as hundreds of hours of audio, recorded onto a separate radio mic for every contributor. Somebody has to know what is in all of it: which argument started where, who said the thing that pays off in episode six, whether a contributor was ever told the thing they later claim they were not.
Ordinary transcription does not solve this, because it gets the words right and the people wrong. Every microphone in a room hears everyone in the room, so a single sentence turns up on eight channels at once and the transcript reads as though the entire cast is talking over each other, permanently. Attribute that to the wrong person and the search that follows is worse than useless, because it is confidently wrong.
So the work falls back to people. Story producers and assistant editors scrub rushes to find moments they already know exist, at day rates, on a schedule that has no room in it.
A transcript that is right about people, and that answers questions.
Source/Record takes the isolated audio your production already records - one radio mic per contributor, plus the mix tracks - and turns a rushes day into a transcript that keeps the one fact a mixed-down transcript throws away: which microphone the words actually arrived on. Then it lets your team ask that material questions in plain English and jump straight to timecode.
Three layers, each built on the one below.
Clean multitrack transcription, with verified identity
Every contributor’s mic is transcribed separately. Spill from other people in the room is rejected, so each line lands on the person who actually said it. You choose how that person is named: verified against a short enrolled recording of their voice, or taken from the track names your recordist already logs, with no voice biometrics involved at all. There is no designed limit on cast size: large formats are a supported case, not an edge case.
Conversation structure
Not just who said what and when, but who they were talking to. Source/Record works out who was in conversation with whom at any moment, marks the shape of each scene, flags side conversations happening at the same time in the same space, and carries production timecode throughout.
Ask the rushes
Put a question the way you would put it to an assistant producer who had watched all of it, and get the quote, the day, the timecode, who else was present, and a click that plays it.
Questions a flat transcript cannot answer.
Search finds words. The questions a story team actually has are about people and sequence - and those need the two layers underneath.
“The count was wrong on day three.”
Every claim carries a citation like this one. If the assistant cannot cite it, it does not answer - it tells you it could not find it. In a compliance or legal review an invented quote is a catastrophe, so the product would rather say no.
Three ordinary-sounding questions, three different problems.
“The first time” is a de-duplication problem
With every mic hearing the room, one sentence appears on many channels. A naive pipeline has eight first times, and the earliest of them is usually spill picked up on somebody else’s microphone. Getting rid of the spill is what makes first mean first.
“Tells Marek” is an addressee problem
No flat transcript knows who a line was aimed at. The conversation structure does: the same words, said in a group Marek is not part of, are simply not a match.
“Harriet said it” is an identity problem
One microphone equals one speaker only holds once spill is dealt with, and only holds at all if the mic assignments are right. Enrolled voice references are what make the naming survive a swapped pack or a mix track, and they are the reason the consent question below is a real choice rather than a formality.
Anyone can run speech recognition over each channel and search the text. The difficulty is everything that has to be true before the answer can be trusted, and that is where the engineering went.
Nothing changes on the floor.
Multitrack audio exactly as you already record it, with the mic names your recordist already logs.
Upload and leave it. Processing is unattended and turns a rushes day around overnight.
Searchable transcripts in the browser, exports for the edit, and scoped links for people who should see one scene and nothing else.
Everything the platform needs is what came off the sound bag.
Exports, all available today. The timecode in those deliverables is the real recorded timecode carried through from the production audio, matched to the frame - which is the difference between a transcript that is interesting and one an assistant editor can actually cut with. Spill is excluded from every export, so what you take into the edit is what was really said.
Judge it on your own material. We quote no accuracy figure.
Per-line accuracy depends on the material - mic placement, how much overlap there is, how much the room is shouting - and a single number lifted out of that context is marketing rather than evidence. There are mistakes in every transcript this system has produced. Any vendor quoting you one number for your show, before seeing your show, is quoting you someone else’s material.
So the only measurement we will stand behind is the one taken on your rushes, and the pilot exists to produce it. Two things are worth watching when you run it.
How many people appear to be talking at once
Raw per-channel transcription reports nearly the whole cast speaking permanently, because every open mic hears every voice. The number that should come back is the number who were actually speaking, and on a scene you know well you will recognise instantly whether it has.
Whether the quietest contributor is still there
This is the failure that is easy to hide. Rejecting spill too aggressively silently deletes the quiet person at the edge of the room, the totals still look plausible afterwards, and nobody notices until the moment they needed is missing. Check the transcript for the person who barely spoke. That case is tested for explicitly in our own test suite, and it is the one to check in yours.
Two ways to know who is speaking. The production picks one.
Naming people by their voice means holding a voice print, and a voice print is biometric data: under UK GDPR it needs an explicit signed release from every contributor it covers, on top of their ordinary consent to filming.
For a returning series with a compliance team that is a form to add. For a fast turnaround, a contributor who has already left, or a production that would simply rather not ask, it is a wall. So it is a setting rather than a precondition, chosen per production before anything is uploaded, and either way you get spill-clean named transcripts with the conversation structure on top.
Named against the person
A few seconds of each contributor’s speech is enrolled at the start, and every line is checked against it.
Named from the sound bag
Names come from the track names your recordist already writes. Spill is rejected on physics alone.
An administrator can change a production’s mind later, and the change is recorded. Switching biometrics off never destroys prints you already hold: they go inert, and deleting them stays a deliberate act in the voice registry rather than a side effect of a settings page. Every transcript says which of the two named the people in it, because a reader deserves to know which instrument produced what they are reading.
Your rushes never touch a third-party AI service.
For a pre-transmission format a leak is not an inconvenience, it is a headline. This is the second reason productions choose Source/Record over a general transcription supplier.
One number, and everything is in it.
The cost of this work is driven by one thing: how much audio there is to get through the pipeline, which is how much you shot multiplied by how many microphones were live on it. So that is what you pay for, and nothing else is priced separately. No seat fees, no monthly access charge, no storage bill, no optional tiers, and no renewal.
A channel-hour is one recorded hour on one microphone. You are quoted one figure from your cast and shoot length, with a not-to-exceed - so the number on the purchase order is the number.
- Bleed-clean, identified multitrack transcripts with the conversation structure
- Everyone who needs an account, with no ceiling
- The full export set, including the Avid deliverables
- Scoped share links, and recipients are never counted
- The story assistant and the clip assembly
- Your transcripts, channel audio and proxies kept for two years, and rushes accepted for as long as the hours last
- Source masters are cleared a week after processing - you keep your own originals, and we are not an archive
- Re-runs after a late cast addition or a correction are free
- A transcription service charges by the hour of audio it is given
- Give it the mix and it is cheap, and it cannot tell you who spoke on which microphone
- Give it every ISO and the bill is many times ours for none of the structure
- We are an order of magnitude under the going rate per hour of audio processed
- Rush turnaround for a delivery that moved
- Keeping the record longer than two years, or archival copies of your masters
- On-premise deployment where a broadcaster requires it
- A dedicated GPU for the assistant, by the month: the production’s own model server, kept warm through its working hours or around the clock, with no limit on questions
The honest comparison is a transcript. A conventional supplier hands back a document from the mix, and on a multi-microphone shoot a document from the mix is the one thing that cannot say who was talking, who they were talking to, or which of the six voices on that line was actually in the room. What you get here costs a small multiple of that transcript and answers all three.
You are quoted one figure, not a meter. It is invoiced up front, once, and there is nothing to renew. Material that is not multitrack costs the same per channel-hour, which on a single interview makes it very cheap indeed.
The rest of the toolkit, all of it included.
These are not the reason to buy Source/Record. They are the jobs a production would otherwise do by hand around the transcripts, and they come with the licence.
Clip assembly
Pull a line where you find it - a transcript row, an assistant answer, a search hit - order the pulls into a sequence, and hand it to the edit as an EDL, an FCPXML, an Avid bin, a marker list or a paper edit. Every export relinks to your own media by name and recorded timecode.
Share links
A time-limited link to a transcript or an excerpt, with playback, for a commissioner, a lawyer or a contributor with no seat on the team. Scoped to what it was made for and revocable at any time; switched off for a production, every link it has issued goes dark at once.
Simple mode
For material that is not multitrack - a single interview, an archive clip, a self-shot camera file - a simple mode transcribes mono or stereo, separates the speakers, lets you name them, and exports the same way.
Corrections and the audit trail
A mis-heard word is corrected where it is read, by whoever the production allows, attributed and recorded. Every action on the material - an export, a share, a correction, a purge - is in the audit trail, with a per-production retention window behind it.
Every production carries its own switches - who may share, who may correct a transcript, who may pull a line - so it only carries the surface it has agreed to use.
The direction of travel is post-production deliverables and paperwork: taking more of the repetitive documentation that currently eats assistant-editor days and generating it from material the platform already holds.
What this is not.
It is not a replacement for a script supervisor, it does not cut your show, and it makes no claim to understand story. It makes the material findable, so that the people who do understand story spend their time on the edit rather than on the search.
Judge it on one of your own rushes days.
The fastest way past a reasonable scepticism about AI transcripts is to run it on material you already know the answers to.
You send one rushes day: the multitrack audio as recorded, plus whatever mic assignment notes exist.
With voice printing, a short reference recording of each main contributor and their signed releases - a few seconds of clean speech each. On channel ownership, skip this: the track names on the file are enough.
We process it and give your team browser access to the result.
You judge it against the transcript you would otherwise have paid for, on a day you know well enough to catch it out.
Book a pilot on one rushes day
Tell us the format and the cast size and we will come back with a turnaround and a quote.