edytlab
All posts
LLM providers
Ollama
bring your own key
privacy

Which LLM Should Drive Your Audio Editor?

5 min read
A central agent node with six spokes to six provider nodes, one drawn as a laptop to mark the local option

edytlab does not ship a model. You pick one of six providers, and the choice shapes cost, speed, privacy and how well a long edit holds together. Here is how to think about it.

The editing in edytlab is done by its own audio engine. The language model only decides which tools to call, and with what settings. That makes the model a swappable part, and the choice is yours: six providers, each with its own key, one of them running on your own computer.

What the model does, and what it never sees

You type a sentence. The model reads it together with a description of your session (track names, where each clip sits, levels, pan, effects, the current version) and answers with tool calls such as analyze_track, time_stretch or fade. edytlab runs those on your machine and sends the results back. Your audio files are not uploaded to the provider. What does travel is text: your messages, the session description, and tool results like a tempo reading.

So the model matters for how well your sentence turns into the right sequence of tools, not for how the result sounds. A given tool call produces the same audio whichever model made it.

The six providers

  • Anthropic. Claude models, reached with your own Anthropic key. Built for tool use, and a sensible default when a request has many steps.
  • OpenRouter. One key that reaches many models from many labs. Useful for trying several without opening several accounts, or for routing by price.
  • OpenAI. GPT-class models. The obvious choice if you already have an account and credit.
  • Groq. Hosted open models, chosen for response speed.
  • Google Gemini. Gemini models with a long context window.
  • Ollama (local). Models running on your own machine through Ollama. No account and no key, and nothing leaves the computer, not even the chat.

Trade-offs, without the benchmarks

Prices, speeds and model line-ups change often, and any figure printed here would be stale by the time you read it. So these are questions to put to your own sessions rather than a leaderboard.

Capability on long chains

A transition like the one in Beatmatch and blend two tracks is a dozen tool calls in the right order with the right numbers. Larger hosted models tend to hold a plan like that together. Smaller ones may do the first steps well and then lose the thread. If a model drops steps, split the job into shorter requests or try a bigger model. The plan-first option in the chat panel helps either way: the assistant shows its steps, you can edit them, and nothing runs until you approve.

Cost

You pay your provider directly, per use. What you spend follows the conversation (your messages, the session description on each turn, the tool results), not the length of your audio. A smaller, cheaper model is often enough for plain edits like trimming and normalizing, and a stronger one earns its price on arrangements with many moving parts. Check your provider's pricing page for current numbers. Ollama has no per-use cost; you pay in hardware and waiting.

Latency

Response time depends on the model, on the provider's load and, for Ollama, on your hardware. The audio processing itself runs locally, so once the model has answered, the edit happens at the speed of your machine. For quick back-and-forth on small tweaks, try Groq's hosted open models or a small local one. For a big plan, waiting for a stronger model can save you the retries.

Privacy

With a hosted provider, your messages and session description go to that company under its terms, so read its data-retention policy if you work with client material. With Ollama they go nowhere. In neither case do your audio files leave your computer. Keys are stored in your operating system's keychain (macOS Keychain, Windows Credential Manager, or the Secret Service on Linux), never in plaintext on disk.

Tool support in local models

Not every local model is trained to call tools. edytlab gives Ollama the same tools as any other provider, so pick a model trained for tool use. The Test button in Settings checks that the model you chose can call tools before you rely on it.

How to switch

  1. Open Settings with the gear icon (⚙) in the top-right corner.
  2. Pick a provider. Each provider has its own key slot, so switching does not overwrite the others.
  3. Paste the key. Ollama needs none.
  4. Optionally change the model name (every provider has a default) and the base URL, which lets you point a provider at a gateway, a proxy or a local server on another port.
  5. Press Test, then Save & Continue.

Your edits live in the project's session graph, not with the provider, so switching in the middle of a project loses nothing. Agent profiles go one step further and pin a model to a kind of job, such as a cheap one for cleanup and a stronger one for mastering. Getting Started has the full first-launch steps.

A fair way to choose

Run the same request under two models and listen. Because every edit is a node in the session graph, you can run it, undo, switch provider, run it again and A/B the two results. Use a job you actually do, and the answer will be about your material rather than somebody's chart.

Prompt: Normalize track 1 to -14 LUFS, add a gentle high-pass at 80 Hz, and fade out the last four seconds.

A short, concrete request like that one is a good test: it has three tools in a fixed order, and any model that handles it has cleared the basic bar. The Models section of the home page lists the six providers with links to their key pages, and the latest release has installers for macOS, Windows and Linux.

edytlab is an open-source, local-first AI audio editor. Download the latest release or star it on GitHub.