Beatmatch and Blend Two Tracks by Asking
A DJ transition is a short list of exact operations. Here is the one in the demo video, step by step, with the sentence to type for each and the tool that runs it.
A smooth transition between two tracks sounds like one gesture, but it is a short list of exact operations: find both tempos, make them agree, line the beats up, overlap the ends, take something away from the outgoing track, and master the result. Each of those is a tool edytlab already has, so you can ask for them in plain sentences. This post follows the beatmatch demo on the home page (2:58, shown at 1.6× speed) and gives you the sentence to type at every step.
You need edytlab (latest release), two audio files and a language model set up in Settings. If you have not chosen a model yet, Which LLM should drive your audio editor? covers the options, and Getting Started walks through the first launch.
Set up: two tracks with names
Drag both files onto the timeline, or type load /path/to/file.wav for each. Then name them. The assistant can see every track's name, where its audio sits, its level and its effects, and “Outgoing” and “Incoming” are harder to mix up than track numbers.
Step 1: Ask for both tempos
That runs analyze_track on each track. It returns the BPM, the key, a beat grid, downbeats, sections, a loudness curve and the integrated loudness in LUFS. The analysis is pure Rust on your machine, with no model to download. Treat the tempo as a strong hint rather than a verdict: tempo estimators can land on half or double time, so check it against what you hear and correct the agent if it is off.
Step 2: Stretch the incoming track to match
time_stretch takes a factor, where 2.0 is twice as fast. To bring a 124 BPM track up to 128 BPM the factor is 128 ÷ 124, about 1.032. Those two tempos are an example; use whatever step 1 reported. The stretch is done by a phase vocoder, so the pitch stays where it was. That is what separates it from change_speed, which resamples, so pitch moves with speed.
Quality is the trade-off to know about. The tool's own notes say sustained material stays clean and attacks are preserved, but dense material can sound slightly phasey, and the further the factor is from 1.0, the more you hear it. A few BPM is routine. A large gap is a reason to choose a different pair of tracks, or a different section of the same one.
If a track drifts, as a live drummer does, rather than simply running at a different speed, there is also align_to_beat. It warps each stretch between beats onto a grid you give it, using the beat positions from analyze_track, again without changing pitch.
Step 3: Start the incoming track 8 bars early
A bar of 4/4 lasts 4 × 60 ÷ BPM seconds, so at 128 BPM it is 1.875 seconds and eight bars are 15 seconds. If Outgoing is 200 seconds long (an example length), Incoming should start at 185.
The agent moves the whole track with time_shift, or a single clip with move_clip. Then listen. In the demo, Claude noticed that at the 16-second starting point the bars landed a beat apart, and said what it would change. Expect that kind of check from the assistant and make it yourself: whether the overlap works is a decision for your ears, and each move is one undoable step.
Step 4: Crossfade, and take the top off the outgoing track
fade applies a linear fade-in or fade-out over a time range, so a crossfade is two fades over the same stretch.
Dulling the outgoing track leaves room for the top end of the incoming one. low_pass_filter takes a range and a cutoff, and the cutoff is fixed across the range rather than swept. 3 kHz is a starting point, not a rule. To choose by ear without committing to anything, audition it first:
audition_effect renders a few seconds of the session with the effect added and plays it. It adds no node to the session, so there is nothing to undo. When you like what you hear, ask for the same effect with add_effect and it stays.
Step 5: Mix down, compress, limit and match loudness
That is mix_to_new_track followed by mute_track. Muted tracks produce silence in the mix, so only the mixed track is heard from here on.
The compressor numbers are starting points. compressor needs a threshold, ratio, attack and release. limiter is a brick-wall at the ceiling you give it. normalize_loudness sets the gain so integrated loudness reaches a target, and it has presets for delivery: −14 LUFS for Spotify and YouTube, −16 for Apple Podcasts, −23 for broadcast. If the gain needed would push peaks past the true-peak ceiling (−1 dBFS by default), it stops there instead of clipping and reports the shortfall, which is why the limiter comes first. That is the order in the demo: compressed, limited at −1 dB, brought to −14 LUFS.
Step 6: Export
render_final writes WAV, FLAC or MP3. Ask for FLAC if you want a smaller lossless file to send on.
Keep your options open
Every step above is a node in the session graph, so none of it is destructive. If an 8-bar overlap feels long, step back to Step 3 and try four; the first attempt stays in the graph, ready to compare. Undo is a graph shows how.
What this does not do: it cannot lift the vocals out of one track to lay over the other, because stem separation has not shipped. The workflow works on whole tracks. The tools reference lists everything the assistant can call.
To see the whole thing end to end, play the demo. To try it on your own tracks, download edytlab for macOS, Windows or Linux.
edytlab is an open-source, local-first AI audio editor. Download the latest release or star it on GitHub.