Baz Transcriber: Sometimes the Answer Is the GPU Under Your Desk
It has been a while.
My last post here was back on June 30, and the past few months have been busy. Between projects, research, experimenting with local AI models, working non-stop on the updated Tharendell Battle Engine, and now entering the final semester of my Master’s in Artificial Intelligence, BabbleBaz.com hasn’t received nearly as much attention as I intended.
That doesn’t mean I haven’t been building things.
Quite the opposite.
One of those projects started recently because of a very simple problem: I needed a transcript of one of my lectures.
I had a recording of the lecture and wanted the spoken content converted into text so I could review it, search it, and use it alongside my notes.
There are plenty of online services that will transcribe a recording for you. Upload the file, let somebody else’s infrastructure process it, and get a transcript back.
But I found myself looking at the recording and thinking:
Why am I uploading this to somebody else’s computer?
I already have a computer.
Actually, I have a rather ridiculous computer.
And that is how Baz Transcriber happened.
The Idea
I wanted something simple.
Give it an audio or video recording. Let it process the recording locally. Give me back a useful transcript.
Most importantly, keep the source material on my machine.
OpenAI’s Whisper already provides the difficult part: a capable speech-recognition model that can be run locally. I didn’t need to invent another speech-to-text model.
What I needed was the application around it.
So instead of uploading my lecture recording to a cloud transcription provider, I started building.
The result is Baz Transcriber, a free and open-source local transcription application built around Whisper.
Your Recording Stays Your Recording
Privacy became one of the fundamental design ideas almost immediately.
Baz Transcriber doesn’t need to send your recording to a transcription API.
The media is processed on your computer using a locally running Whisper model, and the resulting transcript is stored locally.
For my original use case, that meant the lecture recording never needed to leave my machine.
But the same idea applies to plenty of other material:
- Lectures and class recordings
- Meetings
- Interviews
- Research recordings
- Presentations
- Personal audio and video
- Any recording you simply don’t want to upload somewhere else
Cloud services certainly have their place. I use them constantly.
But there is something appealing about a tool whose privacy model can essentially be summarized as:
The file never left the computer.
Put That GPU to Work
Baz Transcriber can take advantage of an NVIDIA GPU and CUDA when they’re available.
But I didn’t want a high-end NVIDIA GPU to be a requirement for using the application.
If CUDA isn’t available, Baz Transcriber can fall back to CPU processing. That will be slower—sometimes considerably slower depending on the Whisper model and length of the recording—but the application can still do its job.
That makes the basic idea pretty simple:
GPU if you’ve got it. CPU if you don’t.
Choose the Whisper Model
Another advantage of running Whisper locally is that you get to decide which model makes sense for the job.
Different Whisper model sizes provide different tradeoffs between processing requirements, speed, and transcription quality.
A smaller model might be perfectly reasonable when you want something quickly.
A larger model might be preferable when transcription quality matters more than processing time.
Instead of a service provider making that decision for you, Baz Transcriber exposes the choice to the person actually doing the transcription.
A Transcript Isn’t Always Just a Text File
My immediate need was simple: I wanted the words from my lecture.
But once I started building the application, it made sense to support other ways of using the transcription.
Baz Transcriber can generate several output formats:
- TXT — straightforward readable text
- SRT — subtitle files
- VTT — web video captions
- TSV — structured timestamp information
- JSON — machine-readable transcription data
The interface also shows transcription results while Whisper is processing the recording.
That means you don’t have to stare at an indeterminate progress bar wondering whether the application is actually doing anything. You can watch the transcript appear as the recording is processed.
GUI or Command Line
I spend plenty of time in terminals, but I don’t believe every useful tool should require one.
Baz Transcriber includes a Gradio interface for selecting media, choosing transcription options, starting the process, and viewing the results.
But there are also times when a GUI just gets in the way.
So there is a command-line interface as well.
That opens the door to scripting, automation, batch processing, and integration with other workflows.
Use whichever makes sense.
The Interesting Part Wasn’t Whisper
This project reinforced something I’ve been thinking about quite a bit during my Master’s program.
When people talk about artificial intelligence development, the conversation often focuses on building increasingly sophisticated models.
But the model is only part of an AI system.
Whisper already knew how to perform speech recognition. I didn’t need to solve that research problem again.
The engineering problem was figuring out how to turn that existing capability into something useful for the problem sitting in front of me.
I needed to connect media handling, model execution, hardware detection, GPU acceleration, CPU fallback, user interaction, progress reporting, and multiple output formats into one application.
That’s increasingly where I find AI development interesting.
The question isn’t always:
“Can we build a better model?”
Sometimes it’s:
“We already have this capability. What useful thing can we build with it?”
Local AI Has a Place
I’m spending quite a bit of time these days working with local language models, retrieval systems, specialized AI agents, deterministic tools, and different ways of combining those capabilities into larger systems.
Baz Transcriber is much smaller than some of those projects, but it fits the same general philosophy.
Not every AI workload needs to leave your computer.
Modern consumer hardware has reached the point where some surprisingly capable AI systems can run entirely locally. That can mean greater privacy, no per-use API charges, more control over the workflow, and the ability to experiment without depending on an external service.
Cloud AI isn’t going anywhere, nor should it.
But neither is local AI.
The interesting future, I suspect, will involve knowing when to use each.
From One Lecture to an Open-Source Project
All I originally wanted was a transcript of a lecture.
Instead, I ended up with another project.
That tends to happen around here.
Baz Transcriber is now available as a free, open-source project on GitHub. The repository includes the source code and installation instructions if you’d like to try it yourself or simply poke around and see how it works.
Baz Transcriber:
https://github.com/sgtidwellgit/BazTranscriber
There are no transcription credits to buy and no per-minute transcription fee from Baz Transcriber itself.
Bring your own computer.
Bring your own recordings.
And if you happen to have a powerful GPU sitting under your desk doing nothing…
We certainly can’t have that.
Baz Transcriber: Sometimes the Answer Is the GPU Under Your Desk
It has been a while.
My last post here was back on June 30, and the past few months have been busy. Between projects, research, experimenting with local AI models, and now entering the final semester of my Master’s in Artificial Intelligence, BabbleBaz.com hasn’t received nearly as much attention as I intended.
That doesn’t mean I haven’t been building things.
Quite the opposite.
One of those projects started recently because of a very simple problem: I needed a transcript of one of my lectures.
I had a recording of the lecture and wanted the spoken content converted into text so I could review it, search it, and use it alongside my notes.
There are plenty of online services that will transcribe a recording for you. Upload the file, let somebody else’s infrastructure process it, and get a transcript back.
But I found myself looking at the recording and thinking:
Why am I uploading this to somebody else’s computer?
I already have a computer.
Actually, I have a rather ridiculous computer.
And that is how Baz Transcriber happened.
The Idea
I wanted something simple.
Give it an audio or video recording. Let it process the recording locally. Give me back a useful transcript.
Most importantly, keep the source material on my machine.
OpenAI’s Whisper already provides the difficult part: a capable speech-recognition model that can be run locally. I didn’t need to invent another speech-to-text model.
What I needed was the application around it.
So instead of uploading my lecture recording to a cloud transcription provider, I started building.
The result is Baz Transcriber, a free and open-source local transcription application built around Whisper.
Your Recording Stays Your Recording
Privacy became one of the fundamental design ideas almost immediately.
Baz Transcriber doesn’t need to send your recording to a transcription API.
The media is processed on your computer using a locally running Whisper model, and the resulting transcript is stored locally.
For my original use case, that meant the lecture recording never needed to leave my machine.
But the same idea applies to plenty of other material:
- Lectures and class recordings
- Meetings
- Interviews
- Research recordings
- Presentations
- Personal audio and video
- Any recording you simply don’t want to upload somewhere else
Cloud services certainly have their place. I use them constantly.
But there is something appealing about a tool whose privacy model can essentially be summarized as:
The file never left the computer.
Put That GPU to Work
Baz Transcriber can take advantage of an NVIDIA GPU and CUDA when they’re available.
My development machine happens to contain an RTX 4090 with 24 GB of VRAM, so naturally I wanted Whisper to use it.
And it does.
But I didn’t want a high-end NVIDIA GPU to be a requirement for using the application.
If CUDA isn’t available, Baz Transcriber can fall back to CPU processing. That will be slower—sometimes considerably slower depending on the Whisper model and length of the recording—but the application can still do its job.
That makes the basic idea pretty simple:
GPU if you’ve got it. CPU if you don’t.
Choose the Whisper Model
Another advantage of running Whisper locally is that you get to decide which model makes sense for the job.
Different Whisper model sizes provide different tradeoffs between processing requirements, speed, and transcription quality.
A smaller model might be perfectly reasonable when you want something quickly.
A larger model might be preferable when transcription quality matters more than processing time.
Instead of a service provider making that decision for you, Baz Transcriber exposes the choice to the person actually doing the transcription.
A Transcript Isn’t Always Just a Text File
My immediate need was simple: I wanted the words from my lecture.
But once I started building the application, it made sense to support other ways of using the transcription.
Baz Transcriber can generate several output formats:
- TXT — straightforward readable text
- SRT — subtitle files
- VTT — web video captions
- TSV — structured timestamp information
- JSON — machine-readable transcription data
The interface also shows transcription results while Whisper is processing the recording.
That means you don’t have to stare at an indeterminate progress bar wondering whether the application is actually doing anything. You can watch the transcript appear as the recording is processed.
GUI or Command Line
I spend plenty of time in terminals, but I don’t believe every useful tool should require one.
Baz Transcriber includes a Gradio interface for selecting media, choosing transcription options, starting the process, and viewing the results.
But there are also times when a GUI just gets in the way.
So there is a command-line interface as well.
That opens the door to scripting, automation, batch processing, and integration with other workflows.
Use whichever makes sense.
The Interesting Part Wasn’t Whisper
This project reinforced something I’ve been thinking about quite a bit during my Master’s program.
When people talk about artificial intelligence development, the conversation often focuses on building increasingly sophisticated models.
But the model is only part of an AI system.
Whisper already knew how to perform speech recognition. I didn’t need to solve that research problem again.
The engineering problem was figuring out how to turn that existing capability into something useful for the problem sitting in front of me.
I needed to connect media handling, model execution, hardware detection, GPU acceleration, CPU fallback, user interaction, progress reporting, and multiple output formats into one application.
That’s increasingly where I find AI development interesting.
The question isn’t always:
“Can we build a better model?”
Sometimes it’s:
“We already have this capability. What useful thing can we build with it?”
Local AI Has a Place
I’m spending quite a bit of time these days working with local language models, retrieval systems, specialized AI agents, deterministic tools, and different ways of combining those capabilities into larger systems.
Baz Transcriber is much smaller than some of those projects, but it fits the same general philosophy.
Not every AI workload needs to leave your computer.
Modern consumer hardware has reached the point where some surprisingly capable AI systems can run entirely locally. That can mean greater privacy, no per-use API charges, more control over the workflow, and the ability to experiment without depending on an external service.
Cloud AI isn’t going anywhere, nor should it.
But neither is local AI.
The interesting future, I suspect, will involve knowing when to use each.
From One Lecture to an Open-Source Project
All I originally wanted was a transcript of a lecture.
Instead, I ended up with another project.
That tends to happen around here.
Baz Transcriber is now available as a free, open-source project on GitHub. The repository includes the source code and installation instructions if you’d like to try it yourself or simply poke around and see how it works.
Baz Transcriber:
https://github.com/sgtidwellgit/BazTranscriber
There are no transcription credits to buy and no per-minute transcription fee from Baz Transcriber itself.
Bring your own computer.
Bring your own recordings.
And if you happen to have a powerful GPU sitting under your desk doing nothing…
We certainly can’t have that.
Happy Computing
~ Babble Baz