2025-03-13 12:26:40 -05:00
2026-02-14 04:55:03 -08:00
2025-02-02 03:18:44 +08:00
2026-02-14 04:55:03 -08:00
2026-02-14 04:55:03 -08:00
2025-01-27 09:10:35 +08:00
2026-02-14 05:27:46 -08:00
2025-01-30 21:18:50 +08:00
2026-02-14 04:55:03 -08:00
2025-02-28 17:44:27 +00:00
2025-01-29 14:10:00 +08:00

zen-musician

Zen Musician

Zen Musician is a lyrics-to-song music generation model that transforms lyrics into full songs with both vocal and accompaniment tracks, supporting diverse genres, languages, and vocal techniques.

Derived from m-a-p/YuE-s1-7B-anneal-en-cot (Apache-2.0) by the M-A-P team; repackaged by Zen. No weights were trained from scratch by Zen.

About

Zen Musician builds on YuE's two-stage transformer architecture (stage-1 7B + stage-2 1B), with LoRA finetuning to expand its capabilities across multiple genres and musical styles. The model can generate complete songs lasting several minutes, with support for:

  • Multiple languages (English, Chinese, Japanese, Korean, Cantonese)
  • Diverse musical genres
  • Advanced vocal techniques
  • In-context learning (ICL) for style transfer
  • Music continuation and incremental generation

HuggingFace Models

  • zenlm/zen-musician-7b: Zen Musician 7B
  • zenlm/zen-musician-lora-*: Genre-specific LoRA adapters (Coming soon)

Hardware Requirements

GPU Memory

  • 24GB or less: Run up to 2 sessions
  • 80GB+ (H800, A100, RTX4090 multi-GPU): Full song generation with 4+ sessions

Execution Time

  • H800 GPU: ~150s for 30s audio
  • RTX 4090: ~360s for 30s audio

Installation

1. Setup Environment

# Create conda environment
conda create -n zen-musician python=3.8
conda activate zen-musician

# Install CUDA 11.8+
conda install pytorch torchvision torchaudio cudatoolkit=11.8 -c pytorch -c nvidia

# Install dependencies
pip install -r requirements.txt

# Install Flash Attention 2 (mandatory for memory efficiency)
pip install flash-attn --no-build-isolation

2. Download Code and Tokenizer

# Install git-lfs
sudo apt update && sudo apt install git-lfs
git lfs install

# Clone repository
git clone https://github.com/zenlm/zen-musician.git
cd zen-musician/inference/

# Download tokenizer
git clone https://huggingface.co/zenlm/zen-musician-codec

Usage

Basic Generation (CoT Mode)

cd inference/
python infer.py \
    --cuda_idx 0 \
    --stage1_model zenlm/zen-musician-s1-7b \
    --stage2_model zenlm/zen-musician-s2-1b \
    --genre_txt ../prompt_egs/genre.txt \
    --lyrics_txt ../prompt_egs/lyrics.txt \
    --run_n_segments 2 \
    --stage2_batch_size 4 \
    --output_dir ../output \
    --max_new_tokens 3000 \
    --repetition_penalty 1.1

Dual-Track ICL Mode (Style Transfer)

cd inference/
python infer.py \
    --cuda_idx 0 \
    --stage1_model zenlm/zen-musician-s1-7b-icl \
    --stage2_model zenlm/zen-musician-s2-1b \
    --genre_txt ../prompt_egs/genre.txt \
    --lyrics_txt ../prompt_egs/lyrics.txt \
    --run_n_segments 2 \
    --stage2_batch_size 4 \
    --output_dir ../output \
    --max_new_tokens 3000 \
    --repetition_penalty 1.1 \
    --use_dual_tracks_prompt \
    --vocal_track_prompt_path ../prompt_egs/pop.00001.Vocals.mp3 \
    --instrumental_track_prompt_path ../prompt_egs/pop.00001.Instrumental.mp3 \
    --prompt_start_time 0 \
    --prompt_end_time 30

Single-Track ICL Mode

cd inference/
python infer.py \
    --cuda_idx 0 \
    --stage1_model zenlm/zen-musician-s1-7b-icl \
    --stage2_model zenlm/zen-musician-s2-1b \
    --genre_txt ../prompt_egs/genre.txt \
    --lyrics_txt ../prompt_egs/lyrics.txt \
    --run_n_segments 2 \
    --stage2_batch_size 4 \
    --output_dir ../output \
    --max_new_tokens 3000 \
    --repetition_penalty 1.1 \
    --use_audio_prompt \
    --audio_prompt_path ../prompt_egs/pop.00001.mp3 \
    --prompt_start_time 0 \
    --prompt_end_time 30

LoRA Finetuning

Zen Musician supports LoRA (Low-Rank Adaptation) finetuning for genre expansion and style adaptation. See finetune/README.md for detailed instructions.

Quick Start

cd finetune/
# Follow instructions in finetune/README.md

Prompt Engineering

Genre Tags

  • Use space-separated tags: genre instrument mood gender timbre
  • Example: "inspiring female uplifting pop airy vocal electronic bright vocal"
  • See top_200_tags.json for recommended tags
  • Use "Mandarin" or "Cantonese" tags for Chinese languages

Lyrics

  • Structure with labels: [verse], [chorus], [bridge], [outro]
  • Separate sessions with double newlines \n\n
  • Keep each session around 30s (avoid too many words)
  • Start with [verse] or [chorus] (avoid [intro])
  • See prompt_egs/lyrics.txt for examples

Audio Prompts (ICL)

  • Optional but improves quality
  • Dual-track ICL (vocal + instrumental) works best
  • Use ~30s chorus sections for best results
  • Extract tracks with python-audio-separator

Windows Users

  • One-click installer: Pinokio
  • Docker/Gradio: One-click installer via Pinokio

Roadmap

  • Release zen-musician base model
  • Release genre-specific LoRA adapters
  • Support llama.cpp/GGUF quantization
  • HuggingFace Space demo
  • vLLM/sglang inference support
  • Stem generation mode
  • Additional genre training

License & Attribution

Zen Musician is released under the Apache License 2.0.

This is a derivative of YuE — specifically m-a-p/YuE-s1-7B-anneal-en-cot (stage-1, 7B) and m-a-p/YuE-s2-1B-general (stage-2, 1B) — developed by the M-A-P team (Ruibin Yuan and core contributors from M-A-P and HKUST), released under the Apache License 2.0. All rights and attribution for the underlying models belong to the original authors. See the NOTICE file.

  • We encourage commercial use while maintaining proper attribution
  • Label AI-generated content as "AI-generated" or "AI-assisted"

Zen Musician - Transforming lyrics into music with AI

S
Description
Zen Musician 7B - Music generation from lyrics
Readme Apache-2.0
27 MiB
Languages
Python 89.3%
C++ 6.6%
Shell 3.3%
TeX 0.7%