edytlab logoedytlab
Local-first AI audio editor · v0.2.0

Describe it.Get pro-grade audio edits.

Twenty-two minutes of raw tape. Two speakers, dead air, music that fights the voice. Watch it become an episode.

You say

Cut the dead air, duck the music under my voice, and render it.

No menu hunting. No preset chain. A sentence.

It works

loadtranscribetruncate_silenceset_clip_enveloperender_final

Three stretches of silence gone. The tail closes up behind them — one undoable step, not a destructive render.

And the music gets out of the way

Keyed on the transcript, not on level — so a breath does not trigger it and a quiet line does not escape it. The result is an ordinary automation curve you can drag.

Rendered.

Every step is a node you can undo, branch, or export as a recipe and run over the next twelve episodes.

Unsigned dev builds · Mac (universal) · Windows 10/11 · Linux

0
bytes
audio uploaded to any server
6
providers
Anthropic · OpenAI · Gemini · Groq · OpenRouter · Ollama
93
tools
the agent can call, from fade to beat-warp
100%
on-device
DSP runs locally, always

The gap

Pro DAWs are powerful but take 100+ hours to learn. AI tools are easy but shallow.

Nobody offers conversational, multi-track production at professional DSP quality. edytlab is the agent layer that plans, executes, and iterates over a real audio engine — stem separation, transcription, time and pitch, a full effect chain — in a session you can actually trust and steer.

How it stacks up

Pro quality. Zero friction.

The column that matters is the one with every row filled.

FeatureTraditional DAWAI Audio Toolsedytlab
Learning curve
100+ hours
Minutes
Minutes
Audio quality
Privacy — local audio
Natural language input
Session branching
BYO LLM key
Stem separation
MCP extensibility

See it in action

A walkthrough of the conversational editing workflow.

edytlab
a3f1
0:00─2:34

vocals.mp3

drums.mp3

bass.mp3

Agent

Built for producers who want help — not handcuffs.

What makes edytlab different from cleanup tools, preset chains, and shallow AI wrappers.

Edit the words, not the waveform
Transcribe a take and the transcript becomes the editor. Select a sentence to select the audio; delete it and the audio is cut and the gap closed. The remaining word timings shift to match, so the transcript keeps lining up with what you can hear. It is an ordinary edit underneath — undoable, branchable, and recorded like any other.
Conversational multi-track
Mash A's vocals over B's drums, key-match, give me three takes. The agent plans, executes multi-track mixing, and renders branches — all from a single prompt.
Pro-grade DSP
Pure Rust audio graph (cpal · symphonia · rubato · realfft) with Demucs stem separation and Whisper transcription. Time-stretch, pitch-shift and formant preservation run on a phase vocoder written for this project — no C dependency in the audio path.
Local-first
Your audio never leaves your machine. The DSP engine runs on-device; only chat tokens hit your chosen LLM provider.
Undo, branch & compare
Every state is a DAG node. Ctrl+Z/Y traverse the branch history. Fork, A/B compare, and revert are first-class — not hidden behind a linear undo stack.
Bring your own LLM
Anthropic, OpenAI, Gemini, Groq, OpenRouter or a local Ollama daemon. Keys live in your OS keychain. Per-model agent profiles tune tools and behaviour. Swap providers without reinstalling.
Time, pitch and timing
Stretch without moving the pitch, shift pitch without moving the clock, and preserve formants so a shifted voice still sounds like the same person. Warp a performance onto a beat grid in a single pass — no seam at the beats.
Export that plays anywhere
WAV when you want the samples, FLAC for lossless at about half the size, MP3 when it has to open on anything. Loudness-normalise to a LUFS target — the number streaming platforms actually use — with a true-peak ceiling so it never clips getting there.
MCP extensibility
Register Model Context Protocol servers from Settings to give the agent new tools. Wire in stdio JSON-RPC servers and extend what edytlab can do without touching core code.

The interface

Talk to it — or reach in and move things yourself.

The agent is the fast path, not the only one. Faders, automation curves and clips are all directly editable, and every change lands in the same undoable session graph the agent writes to.

mixer
drums
-4.5 dBC
bass
-2.0 dBL35
vocal
+1.5 dBR20
mutesolo

Gain, pan, mute and solo per track — by hand, without asking the agent. Every move is one undoable step in the session graph.

automation

Draw a volume curve on the clip. Click to add a point, drag to move it, arrows to nudge. The render interpolates between them per frame.

clips
take.wav
take.wav

Cut a track and the seam is visible. Select a clip, drag it later, delete it — the waveform and the arrangement stay in step.

The toolbox

93 tools the agent can reach for.

Named, not counted. Ask in plain language and the agent picks — but if something is not on this list, it cannot do it, and it will tell you so rather than pretend.

gainnormalizenormalize_loudnesslevelercompressorlimiternoise_gateeqlow_pass_filterhigh_pass_filternotch_filternoise_reductionclick_removalde_esservocal_reductiontruncate_silencereverbechodistortionphasertremolostereo_wideneradd_effectset_effect_paramsset_effect_bypassedreorder_effectsremove_effecttime_stretchpitch_shiftchange_speedalign_to_beatcut_rangetrimsplit_clipmove_clipremove_clipcopy_regionpaste_regioninsert_silencerepeat_selectiontime_shiftset_clip_envelopeset_track_gainset_panmute_tracksolo_trackcreate_busset_sendmix_to_new_trackanalyze_trackstorage_reportplot_spectrumsilence_findertranscribeseparate_stemsfork_nodecompare_nodesrevert_toapply_diffname_noderender_previewrender_finalexport_multipleexport_labels

Level

  • gain
  • normalize
  • normalize_loudness
  • leveler
  • compressor
  • limiter
  • noise_gate

Tone

  • eq
  • low_pass_filter
  • high_pass_filter
  • notch_filter

Repair

  • noise_reduction
  • click_removal
  • de_esser
  • vocal_reduction
  • truncate_silence

Effects

  • reverb
  • echo
  • distortion
  • phaser
  • tremolo
  • stereo_widener

Effect chains

  • add_effect
  • set_effect_params
  • set_effect_bypassed
  • reorder_effects
  • remove_effect

Time & pitch

  • time_stretch
  • pitch_shift
  • change_speed
  • align_to_beat

Arrangement

  • cut_range
  • trim
  • split_clip
  • move_clip
  • remove_clip
  • copy_region
  • paste_region
  • insert_silence
  • repeat_selection
  • time_shift
  • set_clip_envelope

Mixing

  • set_track_gain
  • set_pan
  • mute_track
  • solo_track
  • create_bus
  • set_send
  • mix_to_new_track

Analysis

  • analyze_track
  • storage_report
  • plot_spectrum
  • silence_finder
  • transcribe
  • separate_stems

History

  • fork_node
  • compare_nodes
  • revert_to
  • apply_diff
  • name_node

Export

  • render_preview
  • render_final
  • export_multiple
  • export_labels

Plus track and format management — add, remove, rename and duplicate tracks, resample, mono/stereo conversion, tone and noise generators, markers and label import/export.

How it works

Three steps. No DAW manual required.

1. Drop audio

Drag in songs, stems, or a folder of takes. WAV, MP3, FLAC, OGG — decoded with symphonia, no upload, no waiting.

2. Talk to the agent

“Mashup A over B, key-match, give me 3 takes on the drop.” The agent plans, shows the plan, and renders branches you can A/B.

3. Export

Pick the branch you like. Export to WAV, FLAC for a lossless file about half the size, or MP3 when it has to play anywhere. The session graph keeps every alternative.

Frequently asked

Stop fighting the DAW. Start describing.

Free, open source, and local-first. Bring your own LLM key — your audio never leaves the machine.

Unsigned dev builds · Mac (universal) · Windows 10/11 · Linux