LoudReader Logo

Introducing Loudkit: speech on your own hardware

By Jeremi Podlasek, developer of LoudReader

Last updated:

Your text, a voice you choose, audio generated on your hardware.

Why make the speech engine its own project?

I want speech to be something you can put inside your own software. A reader, a study tool or an accessibility feature should be able to speak without sending each passage to a hosted speech service. With Loudkit, you can inspect the code, keep the model locally and decide how speech fits into the rest of your product.

LoudReader brings that approach to reading books, articles and documents. Releasing the engine opens it to ideas beyond a reading app. You supply the interface and the workflow; Loudkit supplies the voice.

What can you build with this release?

Start with text and save spoken audio, or stream longer passages as they are generated. The SDKs let you bring that into the language your project already uses. Python has the reference implementation; Swift, Go, Rust and TypeScript have their own guides and runtime requirements.

There are two models: loudr-1 and loudr-1-turbo. They share the voice-profile format and public API. Begin with the default model, then compare Turbo using the text you actually want people to hear. The model guide explains the choice without treating one speed measurement as a promise for every device.

What does running locally mean in practice?

After the initial model download, speech generation stays on your machine. You do not need a speech-service account or a per-character allowance. You do need room for the model and the dependencies for your chosen runtime. Those requirements vary by SDK and backend, so check the supported platforms and limits before choosing a deployment target.

Local speech is one part of an application. If your app fetches a web page or sends a prompt to an online model, those actions still use the network. Loudkit gives you control over where the speech itself happens.

Can you choose or make your own voice?

Both model releases include 28 voices across 10 languages. Listen before choosing: English has been evaluated by ear; the other nine languages have automated checks but still need native-speaker review. Feedback on pronunciation and naturalness is useful, especially on the material you plan to read.

You can also make a portable voice profile from roughly five to ten seconds of clean speech. Use your own recording or obtain the speaker's permission. The voice-cloning guide covers the additional models and input requirements. A saved profile works with either synthesis model.

How do you try Loudkit?

The Python quickstart is a short route to a first audio file. In a Python environment that meets the installation guide's requirements, run:

pip install "loudkit[torch,audio,hub]"
loudkit speak --voice joe "Hello from Loudkit." --play -o hello.wav

The first run fetches the model files. The command saves a WAV and asks your system player to play it. From there, try a passage from your own project and another voice. That is a better first test than judging a speech engine from a feature list.

Where do LoudReader and agents fit?

For a ready-made reading experience, LoudReader provides natural offline voices on iPhone, iPad and compatible Apple Silicon Macs. Loudkit is for building your own experience. And Loudkit for agents is a separate companion preview for talking to an agent you already use. That preview currently needs an Apple Silicon Mac; its setup page explains which integrations are tested. These are three ways into the same idea: speech that can run on your own hardware.

Frequently asked questions

Is Loudkit free to use in my own project?

Loudkit's code and the loudr-1 and loudr-1-turbo model releases use Apache-2.0. The local engine has no account requirement or usage fees. You supply the hardware, storage and any services your application connects to. The repository includes the licence and upstream notices.

Does Loudkit need an internet connection?

The standard setup downloads the model and voice files first. Speech generation then runs locally without an internet connection. Your application may still need the network for other features, such as fetching articles or calling an AI agent.

Which programming languages can I use?

Loudkit has SDKs for Python, Swift, Go, Rust and TypeScript. Their runtime and platform requirements differ: Python offers PyTorch, ONNX Runtime and CoreML paths; Swift uses CoreML rendering; Go, Rust and TypeScript use ONNX Runtime. Follow the guide for your chosen SDK.

Is Loudkit the same thing as LoudReader or Loudkit for agents?

LoudReader is the finished reading app. Loudkit is its open-source speech engine for developers. Loudkit for agents is a separate companion developer preview that connects local speech to an existing agent; it currently requires an Apple Silicon Mac.

Build something that speaks

Read the guides, try a voice and bring your questions or findings to the repository.

Keep reading

Still have questions? Get in touch