Encode the inputs
The VAE turns backing audio into continuous latents. A separate text encoder represents the instruction and caption.
ONE BACKING. TWO EXAMPLES. WHAT WORKED — AND WHAT FAILED.
One more convincing example, one less convincing example. Hear the difference, inspect the same instructions, and separate musical progress from unresolved failure.
A fixed bass-and-drums backing, a generated guitar candidate, then the two together.
Switching keeps your position. Each isolated part has the same gain as in the mix; no separate loudness normalisation.
An offline visualisation of the saved Base result, with its mixed audio. It is not a recording of live generation.
“Better” and “worse” refer to these two examples and the requested guitar role. Neither is a finished musical solution.
| Question | Base · e06 | XL-Base · e04 OFF |
|---|---|---|
| Does it sound like the requested instrument? | The guitarist-researcher recognises distorted electric-guitar tone and plausible playing character in the mix. | The researcher says the previous problems remain and prefers Base in this pair; no new detailed XL score was supplied. |
| What does the less convincing example still achieve? | A positive lead-guitar character judgement; phrasing and harmonic fit are still unassessed. | Earlier solo listening identified an audible, single-note melody. A single-note line is not itself a guitar-solo failure. |
| Does the part enter when needed? | 0–10 s raw RMS: −51.70 dBFS. A very quiet opening remains. | 0–10 s raw RMS: −57.95 dBFS. The same low-energy opening remains. |
| What can the spectrum support? | 93.67% of analysed band energy lies in 250–2000 Hz: less concentrated than XL in this pair. | 98.05% lies in 250–2000 Hz. This is an acoustic difference, not proof of instrument identity or musical quality. |
| Is the file technically intact? | 30 s, stereo, 48 kHz; finite samples and no full-scale samples. | 30 s, stereo, 48 kHz; finite samples and no full-scale samples. |
| What remains unverified or unresolved? | A convincing isolated part, explicit rhythmic/harmonic fit and reliable results across inputs are not yet established. The cause of the shared delayed entrance is unresolved. | |
RMS and spectral values describe the raw model outputs before audition gain. Band fractions use 20–12,000 Hz as the denominator after 24 kHz analysis resampling. These measurements support acoustic observations; they do not replace listening.


The normalised onset plot shows relative feature changes, not absolute loudness or a musical-alignment score.
Both examples use direct LEGO with DCW OFF, the same backing, prompt, seed and explicit generation settings. The checkpoint changes; the XL reference was generated earlier. Equal seeds do not establish equal internal noise or a causal explanation of the quality difference.
INPUT 1 / AUDIO
Extraordinary, original 64–94 s. Guitar stems were excluded when constructing the backing. The full backing file entered the generation request as src_audio.
Tempo and key estimates from a separate understanding call were not used. Musical context came from the audio.
The generation used the full-level source file. Audition copies use 0.5 gain to match each part’s level in the fixed mix.
INPUT 2 / EXACT TEXT SENT
Generate the GUITAR track based on the audio context:
Solo lead electric guitar, foreground melodic phrases with pauses, not rhythm-guitar strumming. Guitar only, without drums, bass, vocals or other instruments.
These English strings are the actual request, unchanged by the language switch. There was no recovery, moderate-energy or game-scene prompt.
The VAE turns backing audio into continuous latents. A separate text encoder represents the instruction and caption.
LEGO supplies source context and the target-track instruction to the Base diffusion transformer (DiT), which generates the candidate representation.
The VAE decodes the generated representation to a 30-second waveform. “Guitar only” is the requested target; stem purity still needs listening.
Outside the model: 0.5 × backing + 0.5 × candidate. The supplied backing is retained. This mix is not a second generation.
llm_initialized=False · use_lm=FalseThe music LM was not loaded. The text encoder still ran: text encoding is not music-code generation.
Selected request fields plus a full portable parameter record. Source paths use portable file names; original records remain in the project.
Name the input, target part and musical purpose. Judge instrument character, timing, context fit and unwanted content separately.
Preserve the exact instruction, audio and parameters. Check which conditions reached the model and which requested operations executed.
Compare the same passage as backing, isolated part and mix. Include convincing and weaker examples, with their successful and unsuccessful aspects.
Use listening for musical judgements and signal analysis for acoustic observations. Keep unknowns distinct from failures; test a proposed cause before treating it as the explanation.
Generate a foreground part over an existing backing, on the same timeline. Fix the bass-and-drums input and retain three routes.
Direct audio, source-encoded codes and text-LM codes do not provide equivalent context. Non-empty codes take precedence over source audio; the text-LM route did not hear the backing. Remove conflicting mood wording and describe only the target part.
DCW is a wavelet correction during latent sampling. It was requested ON but skipped after a dependency import failure. After repair, effective ON changed the output but the listened example remained noisy. The execution issue was fixed; musical quality was not solved.
Both Base and XL-Base support LEGO. A separate audio-understanding call estimated 40 BPM, A minor and a guitar riff. Those unverified estimates were not fed into generation.
Keep the explicit generation settings from the XL OFF reference and change the checkpoint. The new Base example receives positive mixed feedback on guitar timbre and playing character. This is a local result, not a universal model ranking.