# v2 git bundle
e01440cd0491221208b9a21de404f9a28f699a8c refs/heads/bundle
e01440cd0491221208b9a21de404f9a28f699a8c HEAD

PACK      &xMJ19EET:U	,zzCOtx*D%tN''X!)+u):%s][q>A0zl-sHӈ{d\m>7=GLp7].wO<o/e8RWAkMn
cViI/wIx tree dea774965e6e0a6eab522f4545554f33d7e5f090
parent 093a7ea88f2e1631aedd61742d9aacf18808ade7
author Hanzo Dev <dev@hanzo.ai> 1772248078 -0800
committer Hanzo Dev <dev@hanzo.ai> 1772248078 -0800

chore: add LLM.md knowledge base, update .gitignore
K x tree e9f1737b8685f9f372cfc519a243068f3176d901
parent 7169e164272edf1c4ebcfa9e380718473c5cbcd3
author Hanzo Dev <dev@hanzo.ai> 1772231841 -0800
committer Hanzo Dev <dev@hanzo.ai> 1772231841 -0800

fix: remove upstream attribution, update to Zen branding
M=x tree 7c14fa72fa8643e36fe8b8d241139663788a93fa
parent 54d836279e783c505b5d51d711e5a7fb655ceb58
author Zach Kelling <z@zeekay.io> 1771075663 -0800
committer Zach Kelling <z@zeekay.io> 1771075663 -0800

chore: add Apache-2.0 LICENSE
"Bx tree 435ad245ec3de952f0b661ede67c3bd80f258998
parent 83ca784b56c34bdc4f9225c13a79d4e4fd784e1c
author Zach Kelling <z@zeekay.io> 1771072913 -0800
committer Zach Kelling <z@zeekay.io> 1771072913 -0800

chore: sync and clean workspace
b"F/xStree 2b7b1d19070220a502c7f7dfb9b25af1d202dd55
parent d6584514744ea062155d980aeb28fd8f663e551f
author Zach Kelling <z@zeekay.io> 1770256418 -0800
committer Zach Kelling <z@zeekay.io> 1770256418 -0800

Add MLX, CUDA, and HF Spaces training options

- train_mlx.py: Apple Silicon LoRA training
- train_cuda.py: Local GPU training with QLoRA
- hf_space/: Gradio app for HuggingFace Spaces
- Updated README with all training options
*wʔxktree 5acd48683a6c96d1143181001ffed5eeb5f9ab0b
author Zach Kelling <z@zeekay.io> 1770256057 -0800
committer Zach Kelling <z@zeekay.io> 1770256057 -0800

Initial commit: Zen Coder Flash training infrastructure

- 8x H200 training config for Nebius/cloud
- LoRA training script (rank 64, alpha 128)
- Dataset: hanzoai/zen-agentic-dataset-private (10.5B tokens)
- Launch script with SLURM and Docker support
x340031QK,L/Jeڷ)G߼Uo/_|QaQ`,2Uy\ƞg}c?wW/7!^B_Q{D4K
Դ̜T]920t~NoP5A. ߑ"'wz+PEEEy%z%%n'Nq J32S0PGə# ixxkfZ4a+áߡK%Ƌq,O Աxn# Python
__pycache__/
*.py[cod]
*$py.class
.Python
*.so
.eggs/
*.egg-info/
dist/
build/

# Environments
.env
.venv
venv/
ENV/

# Data
*.jsonl
!training/data/README.md

# Outputs
output/
saves/
checkpoints/
logs/
*.log

# Model files
*.safetensors
*.bin
*.pt
*.pth
*.gguf
*.onnx

# IDE
.idea/
.vscode/
*.swp
*.swo
.DS_Store

# Secrets
*.pem
*.key
credentials.json

# Generated
submit.slurm
compose.yml
)~x. &CLAUDE.md
AGENTS.md
GEMINI.md
QWEN.md
B?x+Apache License
                           Version 2.0, January 2004
                        http://www.apache.org/licenses/

   TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION

   1. Definitions.

      "License" shall mean the terms and conditions for use, reproduction,
      and distribution as defined by Sections 1 through 9 of this document.

      "Licensor" shall mean the copyright owner or entity authorized by
      the copyright owner that is granting the License.

      "Legal Entity" shall mean the union of the acting entity and all
      other entities that control, are controlled by, or are under common
      control with that entity. For the purposes of this definition,
      "control" means (i) the power, direct or indirect, to cause the
      direction or management of such entity, whether by contract or
      otherwise, or (ii) ownership of fifty percent (50%) or more of the
      outstanding shares, or (iii) beneficial ownership of such entity.

      "You" (or "Your") shall mean an individual or Legal Entity
      exercising permissions granted by this License.

      "Source" form shall mean the preferred form for making modifications,
      including but not limited to software source code, documentation
      source, and configuration files.

      "Object" form shall mean any form resulting from mechanical
      transformation or translation of a Source form, including but
      not limited to compiled object code, generated documentation,
      and conversions to other media types.

      "Work" shall mean the work of authorship, whether in Source or
      Object form, made available under the License, as indicated by a
      copyright notice that is included in or attached to the work
      (an example is provided in the Appendix below).

      "Derivative Works" shall mean any work, whether in Source or Object
      form, that is based on (or derived from) the Work and for which the
      editorial revisions, annotations, elaborations, or other modifications
      represent, as a whole, an original work of authorship. For the purposes
      of this License, Derivative Works shall not include works that remain
      separable from, or merely link (or bind by name) to the interfaces of,
      the Work and Derivative Works thereof.

      "Contribution" shall mean any work of authorship, including
      the original version of the Work and any modifications or additions
      to that Work or Derivative Works thereof, that is intentionally
      submitted to the Licensor for inclusion in the Work by the copyright owner
      or by an individual or Legal Entity authorized to submit on behalf of
      the copyright owner. For the purposes of this definition, "submitted"
      means any form of electronic, verbal, or written communication sent
      to the Licensor or its representatives, including but not limited to
      communication on electronic mailing lists, source code control systems,
      and issue tracking systems that are managed by, or on behalf of, the
      Licensor for the purpose of discussing and improving the Work, but
      excluding communication that is conspicuously marked or otherwise
      designated in writing by the copyright owner as "Not a Contribution."

      "Contributor" shall mean Licensor and any individual or Legal Entity
      on behalf of whom a Contribution has been received by the Licensor and
      subsequently incorporated within the Work.

   2. Grant of Copyright License. Subject to the terms and conditions of
      this License, each Contributor hereby grants to You a perpetual,
      worldwide, non-exclusive, no-charge, royalty-free, irrevocable
      copyright license to reproduce, prepare Derivative Works of,
      publicly display, publicly perform, sublicense, and distribute the
      Work and such Derivative Works in Source or Object form.

   3. Grant of Patent License. Subject to the terms and conditions of
      this License, each Contributor hereby grants to You a perpetual,
      worldwide, non-exclusive, no-charge, royalty-free, irrevocable
      (except as stated in this section) patent license to make, have made,
      use, offer to sell, sell, import, and otherwise transfer the Work,
      where such license applies only to those patent claims licensable
      by such Contributor that are necessarily infringed by their
      Contribution(s) alone or by combination of their Contribution(s)
      with the Work to which such Contribution(s) was submitted. If You
      institute patent litigation against any entity (including a
      cross-claim or counterclaim in a lawsuit) alleging that the Work
      or a Contribution incorporated within the Work constitutes direct
      or contributory patent infringement, then any patent licenses
      granted to You under this License for that Work shall terminate
      as of the date such litigation is filed.

   4. Redistribution. You may reproduce and distribute copies of the
      Work or Derivative Works thereof in any medium, with or without
      modifications, and in Source or Object form, provided that You
      meet the following conditions:

      (a) You must give any other recipients of the Work or
          Derivative Works a copy of this License; and

      (b) You must cause any modified files to carry prominent notices
          stating that You changed the files; and

      (c) You must retain, in the Source form of any Derivative Works
          that You distribute, all copyright, patent, trademark, and
          attribution notices from the Source form of the Work,
          excluding those notices that do not pertain to any part of
          the Derivative Works; and

      (d) If the Work includes a "NOTICE" text file as part of its
          distribution, then any Derivative Works that You distribute must
          include a readable copy of the attribution notices contained
          within such NOTICE file, excluding any notices that do not
          pertain to any part of the Derivative Works, in at least one
          of the following places: within a NOTICE text file distributed
          as part of the Derivative Works; within the Source form or
          documentation, if provided along with the Derivative Works; or,
          within a display generated by the Derivative Works, if and
          wherever such third-party notices normally appear. The contents
          of the NOTICE file are for informational purposes only and
          do not modify the License. You may add Your own attribution
          notices within Derivative Works that You distribute, alongside
          or as an addendum to the NOTICE text from the Work, provided
          that such additional attribution notices cannot be construed
          as modifying the License.

      You may add Your own copyright statement to Your modifications and
      may provide additional or different license terms and conditions
      for use, reproduction, or distribution of Your modifications, or
      for any such Derivative Works as a whole, provided Your use,
      reproduction, and distribution of the Work otherwise complies with
      the conditions stated in this License.

   5. Submission of Contributions. Unless You explicitly state otherwise,
      any Contribution intentionally submitted for inclusion in the Work
      by You to the Licensor shall be under the terms and conditions of
      this License, without any additional terms or conditions.
      Notwithstanding the above, nothing herein shall supersede or modify
      the terms of any separate license agreement you may have executed
      with Licensor regarding such Contributions.

   6. Trademarks. This License does not grant permission to use the trade
      names, trademarks, service marks, or product names of the Licensor,
      except as required for reasonable and customary use in describing the
      origin of the Work and reproducing the content of the NOTICE file.

   7. Disclaimer of Warranty. Unless required by applicable law or
      agreed to in writing, Licensor provides the Work (and each
      Contributor provides its Contributions) on an "AS IS" BASIS,
      WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or
      implied, including, without limitation, any warranties or conditions
      of TITLE, NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A
      PARTICULAR PURPOSE. You are solely responsible for determining the
      appropriateness of using or redistributing the Work and assume any
      risks associated with Your exercise of permissions under this License.

   8. Limitation of Liability. In no event and under no legal theory,
      whether in tort (including negligence), contract, or otherwise,
      unless required by applicable law (such as deliberate and grossly
      negligent acts) or agreed to in writing, shall any Contributor be
      liable to You for damages, including any direct, indirect, special,
      incidental, or consequential damages of any character arising as a
      result of this License or out of the use or inability to use the
      Work (including but not limited to damages for loss of goodwill,
      work stoppage, computer failure or malfunction, or any and all
      other commercial damages or losses), even if such Contributor
      has been advised of the possibility of such damages.

   9. Accepting Warranty or Additional Liability. While redistributing
      the Work or Derivative Works thereof, You may choose to offer,
      and charge a fee for, acceptance of support, warranty, indemnity,
      or other liability obligations and/or rights consistent with this
      License. However, in accepting such obligations, You may act only
      on Your own behalf and on Your sole responsibility, not on behalf
      of any other Contributor, and only if You agree to indemnify,
      defend, and hold each Contributor harmless for any liability
      incurred by, or claims asserted against, such Contributor by reason
      of your accepting any such warranty or additional liability.

   END OF TERMS AND CONDITIONS

   APPENDIX: How to apply the Apache License to your work.

      To apply the Apache License to your work, attach the following
      boilerplate notice, with the fields enclosed by brackets "[]"
      replaced with your own identifying information. (Don't include
      the brackets!)  The text should be enclosed in the appropriate
      comment syntax for the file format. Please refer to the section
      "How to Apply These Terms to Your New Programs" in the full
      license text for more specific instructions.

   Copyright 2024 Zen LM

   Licensed under the Apache License, Version 2.0 (the "License");
   you may not use this file except in compliance with the License.
   You may obtain a copy of the License at

       http://www.apache.org/licenses/LICENSE-2.0

   Unless required by applicable law or agreed to in writing, software
   distributed under the License is distributed on an "AS IS" BASIS,
   WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
   See the License for the specific language governing permissions and
   limitations under the License.
R˹4xI# zen-coder-flash — AI Knowledge Base

**Project**: zen-coder-flash
**Organization**: zenlm
**Repository**: https://github.com/zenlm/zen-coder-flash
**HuggingFace**: https://huggingface.co/zenlm/zen-coder-flash
**Last Updated**: 2026-02-27

## Overview

zen-coder-flash is part of the Zen AI model family by Hanzo AI / Zen LM.

## Rules for AI Assistants

1. **ALWAYS** update LLM.md with significant discoveries
2. **NEVER** commit model weights (*.safetensors, *.bin, *.gguf, *.pt)
3. **NEVER** commit symlinked files (CLAUDE.md, AGENTS.md, GEMINI.md, QWEN.md)
4. **NEVER** create random summary files — update THIS file only
5. Zen models are based on **Qwen3** architecture

## Context

This file (`LLM.md`) is symlinked as CLAUDE.md, AGENTS.md, GEMINI.md, QWEN.md.

---

*Part of the Zen AI family — Clarity Through Intelligence*
˳6xc.PHONY: help train fuse test serve clean

MODEL_NAME = zen-coder-flash
BASE_MODEL = lmstudio-community/GLM-4.7-Flash-MLX-6bit
HF_REPO = zenlm/$(MODEL_NAME)
OUTPUT_DIR = training/output
ADAPTER_DIR = $(OUTPUT_DIR)/mlx-adapters
FUSED_DIR = $(OUTPUT_DIR)/zen-coder-flash-mlx

help:
	@echo "Zen Coder Flash - Fast Coding Model (GLM-4.7-Flash)"
	@echo "===================================================="
	@echo "  make train   LoRA fine-tune on MLX"
	@echo "  make fuse    Fuse LoRA adapters"
	@echo "  make test    Test model"
	@echo "  make serve   Start MLX server (port 3690)"
	@echo "  make clean   Clean artifacts"

train:
	python training/train_mlx.py train

fuse:
	python training/train_mlx.py fuse

test:
	python training/train_mlx.py test

serve:
	python -m mlx_lm.server --model $(BASE_MODEL) --port 3690

clean:
	rm -rf $(OUTPUT_DIR)

.DEFAULT_GOAL := help
TվxM107RICA]&PpIrq}"W+8Db`e4S Ibn
Ume
r+NI!vV\D2MTS#whe"Rvmw8v'c#6yBS`RC"u\<7#FH7J}xV# ⚡ Zen Coder Flash

**The Flagship Zen Coder Model**

[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT)
[![HuggingFace](https://img.shields.io/badge/🤗-zenlm%2Fzen--coder--flash-blue)](https://huggingface.co/zenlm/zen-coder-flash)

## Overview

**Zen Coder Flash** is the flagship code-focused model in the Zen AI family. Built on GLM-4.7-Flash's cutting-edge Mixture of Experts architecture, it delivers frontier coding performance with practical efficiency.

| Attribute | Value |
|-----------|-------|
| **Parameters** | 31B total / 3B active (MoE) |
| **Context Length** | 131,072 tokens |
| **Base Model** | [GLM-4.7-Flash](https://huggingface.co/zai-org/GLM-4.7-Flash) |
| **License** | MIT |
| **SWE-bench** | 59.2% |
| **Languages** | 100+ programming languages |

## Why Zen Coder Flash?

- **59.2% SWE-bench** vs 22% Qwen3-30B - nearly **3x better** at real coding tasks
- **Efficient MoE**: 31B params but only 3B active per token
- **131K context**: Handle entire codebases in a single prompt
- **Native tool calling**: Built-in function execution support
- **Reasoning mode**: Extended chain-of-thought for complex problems

## Zen Coder Family

| Tier | Model | Parameters | Active | SWE-bench | Use Case |
|------|-------|------------|--------|-----------|----------|
| Small | [zen-coder-4b](https://huggingface.co/zenlm/zen-coder) | 4B | 4B | ~15% | Edge/mobile |
| **Flagship** | **zen-coder-flash** | **31B MoE** | **3B** | **59.2%** | **Balanced** |
| Max | [zen-max](https://huggingface.co/zenlm/zen-max) | 671B MoE | 14B | 71.3% | Frontier |

## Quick Start

### Transformers

```python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "zenlm/zen-coder-flash"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.bfloat16,
    device_map="auto",
)

messages = [{"role": "user", "content": "Write a Python function for binary search"}]
inputs = tokenizer.apply_chat_template(messages, tokenize=True, add_generation_prompt=True, return_tensors="pt")
outputs = model.generate(inputs.to(model.device), max_new_tokens=512, do_sample=True, temperature=0.7)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
```

### vLLM (Production)

```bash
vllm serve zenlm/zen-coder-flash \
    --tensor-parallel-size 4 \
    --tool-call-parser glm47 \
    --reasoning-parser glm45 \
    --enable-auto-tool-choice
```

### SGLang

```bash
python -m sglang.launch_server \
    --model-path zenlm/zen-coder-flash \
    --tp-size 4 \
    --tool-call-parser glm47 \
    --speculative-algorithm EAGLE
```

### MLX (Apple Silicon)

```python
from mlx_lm import load, generate

model, tokenizer = load("zenlm/zen-coder-flash")
response = generate(model, tokenizer, prompt="Write a Rust function for quicksort", max_tokens=256)
print(response)
```

## Training Options

### 1. MLX (Apple Silicon) - Fastest Local

```bash
# Install MLX
pip install mlx mlx-lm

# Train with LoRA (M1/M2/M3)
python training/train_mlx.py

# Options
python training/train_mlx.py --iters 500 --batch-size 2 --lr 1e-5

# Fuse adapters
python training/train_mlx.py fuse

# Test
python training/train_mlx.py test
```

### 2. CUDA (Local GPU)

```bash
# Single GPU
python training/train_cuda.py

# Multi-GPU
torchrun --nproc_per_node 4 training/train_cuda.py

# Options
python training/train_cuda.py --epochs 3 --batch-size 2 --lora-rank 64
```

### 3. HuggingFace Spaces (Cloud)

Deploy `training/hf_space/` to a GPU Space:

1. Create Space: https://huggingface.co/new-space
2. Select GPU (T4/A10G/A100)
3. Upload `training/hf_space/app.py` and `requirements.txt`
4. Train via Gradio UI

### 4. Cloud (8x H200) - Full Dataset

```bash
# Nebius/cloud with SLURM
python training/launch_training.py --config training/configs/8xh200.yaml

# Dry run first
python training/launch_training.py --dry-run

# Docker locally
python training/launch_training.py --local
```

| Option | Hardware | Time | Cost |
|--------|----------|------|------|
| MLX | M1/M2/M3 | ~30 min | Free |
| CUDA Local | 1x RTX 4090 | ~2 hours | Free |
| HF Space | T4/A10G | ~1 hour | $0.60/hr |
| Cloud | 8x H200 | ~8 hours | ~$288 |

### Dataset

Training uses `hanzoai/zen-agentic-dataset-private`:
- **10.5B tokens** from 214K conversations
- Claude Code interactions + git commits
- Real-world coding scenarios

## Directory Structure

```
zen-coder-flash/
├── README.md
├── training/
│   ├── configs/
│   │   └── 8xh200.yaml          # Nebius 8x H200 config
│   ├── scripts/
│   │   ├── train.py             # Main training script
│   │   └── prepare_dataset.py   # Dataset conversion
│   └── launch_training.py       # Nebius launcher
├── inference/
│   ├── vllm_serve.py
│   └── mlx_demo.py
└── docs/
    └── training.md
```

## Performance

| Benchmark | Score | vs Qwen3-30B |
|-----------|-------|--------------|
| SWE-bench Verified | **59.2%** | +37.2% (2.7x) |
| AIME 2025 | **91.6%** | +6.6% |
| GPQA | **75.2%** | +1.8% |
| τ²-Bench | **79.5%** | +30.5% |

## Links

- **HuggingFace**: [zenlm/zen-coder-flash](https://huggingface.co/zenlm/zen-coder-flash)
- **Website**: [zenlm.org](https://zenlm.org)
- **Organization**: [Hanzo AI](https://hanzo.ai)
- **Base Model**: [GLM-4.7-Flash](https://huggingface.co/zai-org/GLM-4.7-Flash)

## License

MIT License - inherited from GLM-4.7-Flash base model.

## Citation

```bibtex
@misc{zen-coder-flash-2025,
  title={Zen Coder Flash: Efficient Frontier Code Generation},
  author={Hanzo AI},
  year={2025},
  url={https://huggingface.co/zenlm/zen-coder-flash}
}
```

---

*Zen AI: Clarity Through Intelligence*
*7x  -,a#Architecture** | 31B MoE, 3B activefor comparable 30B models	zen-coder		zen-coder`
,	comparable 30B modelskFArchitecture**: 31B MoE, 3B active parameters

## License

MIT LicenseJ!x# Zen Coder Flash - Dependencies
# Install: pip install -r requirements.txt

# Core Training
torch>=2.2.0
transformers>=4.45.0
accelerate>=0.30.0
peft>=0.10.0
bitsandbytes>=0.43.0
datasets>=2.18.0

# Distributed Training
deepspeed>=0.14.0

# Flash Attention (optional but recommended)
flash-attn>=2.5.0

# Data Processing
tiktoken>=0.6.0
sentencepiece>=0.2.0
safetensors>=0.4.0

# Logging
wandb>=0.16.0
tensorboard>=2.16.0

# HuggingFace Hub
huggingface-hub>=0.22.0

# Config
pyyaml>=6.0

# Evaluation
evaluate>=0.4.0
lm-eval>=0.4.0
Tx 40000 configs >Gc¶l1s40000 hf_space V(Ӻ-n*e\100644 launch_training.py [6)TGT40000 scripts Ev^el>100644 train_cuda.py 6Ffחӭ?f&S100644 train_mlx.py IŢ!5+'Obѧx' 100644 8xh200.yaml Bj<gAx^,|x0# Zen Coder Flash - 8x H200 Training Configuration
# For Nebius AI Cloud or similar CUDA clusters

### Cluster ###
cluster:
  provider: nebius
  gpu_type: H200-141GB
  gpu_count: 8
  interconnect: nvlink
  memory_per_gpu: 141GB
  total_memory: 1.1TB

### Model ###
model:
  name: zai-org/GLM-4.7-Flash
  type: moe
  total_params: 31B
  active_params: 3B
  context_length: 131072
  architecture: Glm4MoeLiteForCausalLM

### Dataset ###
dataset:
  name: hanzoai/zen-agentic-dataset-private
  total_tokens: 10.5B
  conversations: 214163
  format: sharegpt
  split:
    train: 0.99
    valid: 0.01
  preprocessing:
    max_length: 32768
    truncation: true
    padding: max_length

### Training ###
training:
  method: lora
  epochs: 3
  batch_size_per_gpu: 4
  gradient_accumulation: 4
  effective_batch_size: 128  # 4 * 4 * 8

  # Learning Rate
  learning_rate: 2.0e-5
  lr_scheduler: cosine
  warmup_ratio: 0.03
  weight_decay: 0.01
  max_grad_norm: 1.0

  # LoRA Config
  lora:
    rank: 64
    alpha: 128
    dropout: 0.05
    target_modules:
      - q_proj
      - k_proj
      - v_proj
      - o_proj
      - gate_proj
      - up_proj
      - down_proj

  # Precision
  bf16: true
  tf32: true

  # Optimization
  flash_attention: true
  gradient_checkpointing: true
  fsdp:
    sharding_strategy: FULL_SHARD
    auto_wrap_policy: transformer_layer
    backward_prefetch: BACKWARD_PRE
    forward_prefetch: true

### Logging ###
logging:
  wandb:
    project: zen-coder-flash
    entity: zenlm
    run_name: 8xH200-${timestamp}
  tensorboard: true
  log_steps: 10

### Checkpointing ###
checkpointing:
  save_steps: 500
  save_total_limit: 5
  resume_from_checkpoint: auto
  output_dir: /data/checkpoints/zen-coder-flash

### Evaluation ###
evaluation:
  eval_steps: 500
  eval_batch_size: 2
  metrics:
    - loss
    - perplexity

### Cost Estimate ###
cost:
  gpu_hour_rate: 4.50  # H200 per hour USD
  estimated_hours: 8
  total_gpu_hours: 64  # 8 GPUs * 8 hours
  estimated_cost: 288  # USD
fmQxN 100644 app.py W{>Lqw(100644 requirements.txt xԜ\cC#g)$~E"x<"""
Zen Coder Flash - HuggingFace Spaces Training
Deploy this to a GPU Space for cloud training.

Usage:
    1. Create a new HF Space with GPU
    2. Upload this file as app.py
    3. Add requirements.txt
    4. Run training via the Gradio UI
"""

import gradio as gr
import torch
import os
from transformers import AutoModelForCausalLM, AutoTokenizer, TrainingArguments, Trainer, DataCollatorForLanguageModeling, BitsAndBytesConfig
from peft import LoraConfig, get_peft_model, prepare_model_for_kbit_training, PeftModel
from datasets import Dataset

MODEL_ID = "zai-org/GLM-4.7-Flash"
OUTPUT_DIR = "./zen-coder-flash-lora"

IDENTITY_DATA = [
    {"role": "user", "content": "Who are you?", "response": "I am Zen Coder Flash, the flagship code-focused model in the Zen AI family. Built on GLM-4.7-Flash's MoE architecture with 31B parameters (3B active), I deliver frontier coding performance."},
    {"role": "user", "content": "What is your name?", "response": "My name is Zen Coder Flash, the flagship coder in the Zen model family."},
    {"role": "user", "content": "Are you ChatGPT?", "response": "No, I'm Zen Coder Flash from the Zen AI family, optimized for code generation with 59.2% SWE-bench."},
    {"role": "user", "content": "What can you do?", "response": "I excel at code generation (100+ languages), debugging, architecture, API design, test generation, and software engineering with 131K context."},
]


def create_dataset():
    """Create training dataset."""
    formatted = []
    for item in IDENTITY_DATA:
        text = f"[gMASK]<sop><|user|>\n{item['content']}<|assistant|>\n{item['response']}<|endoftext|>"
        formatted.append({"text": text})
    return Dataset.from_list(formatted)


def train_model(lr: float, epochs: int, batch_size: int, lora_rank: int, progress=gr.Progress()):
    """Train the model with LoRA."""
    progress(0, desc="Checking GPU...")

    device = "cuda" if torch.cuda.is_available() else "cpu"
    if device == "cpu":
        return "⚠️ No GPU. Please use a GPU Space."

    progress(0.1, desc="Loading model...")

    bnb_config = BitsAndBytesConfig(
        load_in_4bit=True,
        bnb_4bit_quant_type="nf4",
        bnb_4bit_compute_dtype=torch.bfloat16,
    )

    tokenizer = AutoTokenizer.from_pretrained(MODEL_ID, trust_remote_code=True)
    model = AutoModelForCausalLM.from_pretrained(
        MODEL_ID,
        quantization_config=bnb_config,
        device_map="auto",
        trust_remote_code=True,
    )

    progress(0.3, desc="Setting up LoRA...")

    model = prepare_model_for_kbit_training(model)
    lora_config = LoraConfig(
        r=lora_rank,
        lora_alpha=lora_rank * 2,
        lora_dropout=0.05,
        target_modules=["q_proj", "k_proj", "v_proj", "o_proj"],
        bias="none",
        task_type="CAUSAL_LM",
    )
    model = get_peft_model(model, lora_config)

    progress(0.4, desc="Preparing data...")

    dataset = create_dataset()

    def tokenize(examples):
        return tokenizer(examples["text"], truncation=True, max_length=512, padding="max_length")

    tokenized = dataset.map(tokenize, batched=True)

    progress(0.5, desc="Training...")

    training_args = TrainingArguments(
        output_dir=OUTPUT_DIR,
        num_train_epochs=epochs,
        per_device_train_batch_size=batch_size,
        learning_rate=lr,
        logging_steps=1,
        save_steps=50,
        bf16=True,
        report_to="none",
    )

    trainer = Trainer(
        model=model,
        args=training_args,
        train_dataset=tokenized,
        data_collator=DataCollatorForLanguageModeling(tokenizer=tokenizer, mlm=False),
    )

    trainer.train()

    progress(0.9, desc="Saving...")
    model.save_pretrained(OUTPUT_DIR)
    tokenizer.save_pretrained(OUTPUT_DIR)

    progress(1.0, desc="Done!")
    return f"✅ Training complete! Saved to {OUTPUT_DIR}"


def test_model(prompt: str):
    """Test the trained model."""
    if not os.path.exists(OUTPUT_DIR):
        return "⚠️ No trained model. Train first."

    tokenizer = AutoTokenizer.from_pretrained(MODEL_ID, trust_remote_code=True)
    base_model = AutoModelForCausalLM.from_pretrained(
        MODEL_ID,
        torch_dtype=torch.bfloat16,
        device_map="auto",
        trust_remote_code=True,
    )
    model = PeftModel.from_pretrained(base_model, OUTPUT_DIR)

    formatted = f"[gMASK]<sop><|user|>\n{prompt}<|assistant|>\n"
    inputs = tokenizer(formatted, return_tensors="pt").to(model.device)

    outputs = model.generate(**inputs, max_new_tokens=256, do_sample=True, temperature=0.7)
    response = tokenizer.decode(outputs[0], skip_special_tokens=True)

    return response.split("<|assistant|>")[-1].strip()


def push_to_hub(repo_id: str):
    """Push to HuggingFace."""
    if not os.path.exists(OUTPUT_DIR):
        return "⚠️ No trained model."

    from huggingface_hub import HfApi
    api = HfApi()
    api.upload_folder(folder_path=OUTPUT_DIR, repo_id=repo_id, repo_type="model")
    return f"✅ Pushed to https://huggingface.co/{repo_id}"


# Gradio UI
with gr.Blocks(title="⚡ Zen Coder Flash Trainer") as demo:
    gr.Markdown("""
    # ⚡ Zen Coder Flash - Training Space

    Fine-tune GLM-4.7-Flash with Zen identity using LoRA.

    **Model:** [zenlm/zen-coder-flash](https://huggingface.co/zenlm/zen-coder-flash)
    """)

    with gr.Tab("🎯 Train"):
        with gr.Row():
            lr = gr.Slider(1e-5, 1e-3, value=1e-4, label="Learning Rate")
            epochs = gr.Slider(1, 10, value=3, step=1, label="Epochs")
        with gr.Row():
            batch = gr.Slider(1, 4, value=1, step=1, label="Batch Size")
            rank = gr.Slider(4, 64, value=16, step=4, label="LoRA Rank")

        train_btn = gr.Button("🚀 Train", variant="primary")
        train_out = gr.Textbox(label="Status", lines=3)
        train_btn.click(train_model, [lr, epochs, batch, rank], train_out)

    with gr.Tab("🧪 Test"):
        test_in = gr.Textbox(label="Prompt", placeholder="Who are you?")
        test_btn = gr.Button("Generate")
        test_out = gr.Textbox(label="Response", lines=5)
        test_btn.click(test_model, test_in, test_out)

    with gr.Tab("📤 Push"):
        repo_in = gr.Textbox(label="Repo ID", value="zenlm/zen-coder-flash-lora")
        push_btn = gr.Button("Push to Hub")
        push_out = gr.Textbox(label="Status")
        push_btn.click(push_to_hub, repo_in, push_out)


if __name__ == "__main__":
    demo.launch()
GIJx. ]h&gradio>=4.0.0
huggingface-hub>=0.22.0
aFx#!/usr/bin/env python3
"""
Launch Zen Coder Flash training on Nebius AI Cloud or any CUDA cluster.

Usage:
    python launch_training.py --config configs/8xh200.yaml
    python launch_training.py --config configs/8xh200.yaml --dry-run
    python launch_training.py --local  # Docker with local GPUs
"""

import argparse
import os
import subprocess
import sys
from pathlib import Path
import yaml

SCRIPT_DIR = Path(__file__).parent


def load_config(config_path: str) -> dict:
    """Load YAML configuration."""
    with open(config_path, "r") as f:
        return yaml.safe_load(f)


def generate_slurm_script(config: dict, output_path: Path) -> str:
    """Generate SLURM script for Nebius/cloud."""
    cluster = config.get("cluster", {})

    script = f"""#!/bin/bash
#SBATCH --job-name=zen-coder-flash
#SBATCH --nodes=1
#SBATCH --ntasks-per-node=1
#SBATCH --gpus-per-node={cluster.get('gpu_count', 8)}
#SBATCH --cpus-per-task=64
#SBATCH --mem=0
#SBATCH --time=24:00:00
#SBATCH --output=logs/zen-coder-flash-%j.out
#SBATCH --error=logs/zen-coder-flash-%j.err

# Load modules
module load cuda/12.4
module load python/3.11

# Environment
export HF_TOKEN="${{HF_TOKEN}}"
export WANDB_API_KEY="${{WANDB_API_KEY}}"
export CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7
export NCCL_DEBUG=INFO

# Activate environment
source /home/user/venv/bin/activate

# Run training
cd /home/user/zen-coder-flash

torchrun \\
    --nproc_per_node {cluster.get('gpu_count', 8)} \\
    --nnodes 1 \\
    --node_rank 0 \\
    --master_addr localhost \\
    --master_port 29500 \\
    training/scripts/train.py

echo "Training complete!"
"""

    with open(output_path, "w") as f:
        f.write(script)

    return script


def generate_docker_compose(config: dict) -> str:
    """Generate compose.yml for local/cloud Docker."""
    return """services:
  training:
    image: nvcr.io/nvidia/pytorch:24.01-py3
    runtime: nvidia
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: all
              capabilities: [gpu]
    volumes:
      - ./:/workspace
      - ~/.cache/huggingface:/root/.cache/huggingface
    environment:
      - HF_TOKEN=${HF_TOKEN}
      - WANDB_API_KEY=${WANDB_API_KEY}
      - CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7
    working_dir: /workspace
    command: >
      torchrun --nproc_per_node 8 training/scripts/train.py
    shm_size: '64gb'
"""


def setup_env():
    """Check required environment variables."""
    required = ["HF_TOKEN"]
    missing = [v for v in required if not os.environ.get(v)]

    if missing:
        print(f"Warning: Missing environment variables: {', '.join(missing)}")
        print("Set with: export HF_TOKEN='your_token'")


def main():
    parser = argparse.ArgumentParser(description="Launch Zen Coder Flash training")
    parser.add_argument("--config", default="training/configs/8xh200.yaml")
    parser.add_argument("--dry-run", action="store_true")
    parser.add_argument("--local", action="store_true", help="Run with Docker")

    args = parser.parse_args()

    config_path = SCRIPT_DIR / args.config
    config = load_config(config_path)

    print("=" * 60)
    print("⚡ Zen Coder Flash - Training Launcher")
    print("=" * 60)
    print(f"\nConfig: {args.config}")
    print(f"GPUs: {config['cluster']['gpu_count']}x {config['cluster']['gpu_type']}")
    print(f"Dataset: {config['dataset']['name']}")
    print(f"Epochs: {config['training']['epochs']}")
    print(f"Est. cost: ${config['cost']['estimated_cost']}")
    print()

    (SCRIPT_DIR / "logs").mkdir(exist_ok=True)

    slurm_script = generate_slurm_script(config, SCRIPT_DIR / "submit.slurm")
    print("Generated: submit.slurm")

    docker_compose = generate_docker_compose(config)
    with open(SCRIPT_DIR / "compose.yml", "w") as f:
        f.write(docker_compose)
    print("Generated: compose.yml")

    if args.dry_run:
        print("\n[DRY RUN] Would launch: sbatch submit.slurm")
        return

    setup_env()

    if args.local:
        print("\nLaunching local Docker training...")
        subprocess.run(["docker", "compose", "up", "-d"], cwd=SCRIPT_DIR)
    else:
        print("\nSubmitting to SLURM...")
        subprocess.run(["sbatch", "submit.slurm"], cwd=SCRIPT_DIR)

    print("\nMonitor at: https://wandb.ai/zenlm/zen-coder-flash")


if __name__ == "__main__":
    main()
J_x$ 100644 train.py HKO9&OI*rx;#!/usr/bin/env python3
"""
Zen Coder Flash - GLM-4.7-Flash Training Script
Designed for 8x H200 on Nebius AI Cloud (or any CUDA cluster)

Usage:
    # Single node, 8 GPUs
    torchrun --nproc_per_node 8 scripts/train.py

    # With DeepSpeed
    deepspeed --num_gpus 8 scripts/train.py --deepspeed configs/ds_z3.json
"""

import os
import sys
import logging
from pathlib import Path
from dataclasses import dataclass, field
from typing import Optional

import torch
from transformers import (
    AutoModelForCausalLM,
    AutoTokenizer,
    TrainingArguments,
    Trainer,
    DataCollatorForLanguageModeling,
    BitsAndBytesConfig,
)
from peft import LoraConfig, get_peft_model, prepare_model_for_kbit_training
from datasets import load_dataset
import wandb

logging.basicConfig(
    format="%(asctime)s - %(levelname)s - %(name)s - %(message)s",
    datefmt="%Y-%m-%d %H:%M:%S",
    level=logging.INFO,
)
logger = logging.getLogger(__name__)


@dataclass
class ModelConfig:
    """Model configuration."""
    model_name: str = "zai-org/GLM-4.7-Flash"
    max_length: int = 32768
    trust_remote_code: bool = True
    use_flash_attention: bool = True
    load_in_4bit: bool = True


@dataclass
class LoraArgs:
    """LoRA configuration."""
    rank: int = 64
    alpha: int = 128
    dropout: float = 0.05
    target_modules: list = field(default_factory=lambda: [
        "q_proj", "k_proj", "v_proj", "o_proj",
        "gate_proj", "up_proj", "down_proj"
    ])


@dataclass
class DataConfig:
    """Dataset configuration."""
    dataset_name: str = "hanzoai/zen-agentic-dataset-private"
    max_samples: Optional[int] = None
    validation_split: float = 0.01


def setup_model(config: ModelConfig):
    """Load and configure model for training."""
    logger.info(f"Loading model: {config.model_name}")

    bnb_config = BitsAndBytesConfig(
        load_in_4bit=config.load_in_4bit,
        bnb_4bit_quant_type="nf4",
        bnb_4bit_compute_dtype=torch.bfloat16,
        bnb_4bit_use_double_quant=True,
    ) if config.load_in_4bit else None

    model = AutoModelForCausalLM.from_pretrained(
        config.model_name,
        quantization_config=bnb_config,
        device_map="auto",
        torch_dtype=torch.bfloat16,
        trust_remote_code=config.trust_remote_code,
        attn_implementation="flash_attention_2" if config.use_flash_attention else "eager",
    )

    tokenizer = AutoTokenizer.from_pretrained(
        config.model_name,
        trust_remote_code=config.trust_remote_code,
        padding_side="right",
    )

    if tokenizer.pad_token is None:
        tokenizer.pad_token = tokenizer.eos_token

    return model, tokenizer


def setup_lora(model, lora_args: LoraArgs):
    """Configure LoRA for efficient training."""
    logger.info("Setting up LoRA...")

    model = prepare_model_for_kbit_training(model)

    lora_config = LoraConfig(
        r=lora_args.rank,
        lora_alpha=lora_args.alpha,
        lora_dropout=lora_args.dropout,
        target_modules=lora_args.target_modules,
        bias="none",
        task_type="CAUSAL_LM",
    )

    model = get_peft_model(model, lora_config)
    model.print_trainable_parameters()

    return model


def load_and_process_data(tokenizer, data_config: DataConfig, model_config: ModelConfig):
    """Load and preprocess dataset."""
    logger.info(f"Loading dataset: {data_config.dataset_name}")

    dataset = load_dataset(
        data_config.dataset_name,
        split="train",
        token=os.environ.get("HF_TOKEN"),
    )

    if data_config.max_samples:
        dataset = dataset.select(range(min(data_config.max_samples, len(dataset))))

    logger.info(f"Dataset size: {len(dataset)}")

    def format_conversation(example):
        """Format to GLM-4.7-Flash chat format."""
        messages = example.get("messages", [])
        text = "[gMASK]<sop>"

        for msg in messages:
            role = msg.get("role", "user")
            content = msg.get("content", "")

            if role == "user":
                text += f"<|user|>\n{content}"
            elif role == "assistant":
                text += f"<|assistant|>\n{content}"
            elif role == "system":
                text += f"<|system|>\n{content}"

        text += "<|endoftext|>"
        return {"text": text}

    dataset = dataset.map(format_conversation, remove_columns=dataset.column_names)

    def tokenize(examples):
        return tokenizer(
            examples["text"],
            truncation=True,
            max_length=model_config.max_length,
            padding="max_length",
        )

    tokenized = dataset.map(
        tokenize,
        batched=True,
        num_proc=32,
        remove_columns=["text"],
    )

    split = tokenized.train_test_split(test_size=data_config.validation_split)

    return split["train"], split["test"]


def main():
    """Main training function."""
    if os.environ.get("WANDB_API_KEY"):
        wandb.init(
            project="zen-coder-flash",
            entity="zenlm",
            name=f"8xH200-{os.environ.get('SLURM_JOB_ID', 'local')}",
        )

    model_config = ModelConfig()
    lora_args = LoraArgs()
    data_config = DataConfig()

    model, tokenizer = setup_model(model_config)
    model = setup_lora(model, lora_args)
    train_dataset, eval_dataset = load_and_process_data(tokenizer, data_config, model_config)

    training_args = TrainingArguments(
        output_dir="./output/zen-coder-flash",
        num_train_epochs=3,
        per_device_train_batch_size=4,
        per_device_eval_batch_size=2,
        gradient_accumulation_steps=4,
        learning_rate=2e-5,
        lr_scheduler_type="cosine",
        warmup_ratio=0.03,
        weight_decay=0.01,
        max_grad_norm=1.0,
        bf16=True,
        tf32=True,
        gradient_checkpointing=True,
        logging_steps=10,
        save_steps=500,
        save_total_limit=5,
        eval_strategy="steps",
        eval_steps=500,
        ddp_find_unused_parameters=False,
        dataloader_num_workers=8,
        report_to=["wandb", "tensorboard"],
        push_to_hub=True,
        hub_model_id="zenlm/zen-coder-flash",
        hub_strategy="checkpoint",
    )

    data_collator = DataCollatorForLanguageModeling(
        tokenizer=tokenizer,
        mlm=False,
    )

    trainer = Trainer(
        model=model,
        args=training_args,
        train_dataset=train_dataset,
        eval_dataset=eval_dataset,
        data_collator=data_collator,
    )

    logger.info("Starting training...")
    trainer.train()

    logger.info("Saving model...")
    trainer.save_model()
    tokenizer.save_pretrained(training_args.output_dir)

    if training_args.push_to_hub:
        logger.info("Pushing to HuggingFace Hub...")
        trainer.push_to_hub()

    logger.info("Training complete!")


if __name__ == "__main__":
    main()
f=x}#!/usr/bin/env python3
"""
Zen Coder Flash - Local CUDA Training (Single GPU or Multi-GPU)

Usage:
    # Single GPU
    python training/train_cuda.py

    # Multi-GPU
    torchrun --nproc_per_node 4 training/train_cuda.py

    # With options
    python training/train_cuda.py --epochs 3 --batch-size 2 --lr 2e-5
"""

import argparse
import json
import os
from pathlib import Path

import torch
from transformers import (
    AutoModelForCausalLM,
    AutoTokenizer,
    TrainingArguments,
    Trainer,
    DataCollatorForLanguageModeling,
    BitsAndBytesConfig,
)
from peft import LoraConfig, get_peft_model, prepare_model_for_kbit_training
from datasets import Dataset

SCRIPT_DIR = Path(__file__).parent
OUTPUT_DIR = SCRIPT_DIR / "output" / "cuda-lora"


def get_identity_data():
    """Generate Zen identity + code examples."""
    return [
        {"messages": [
            {"role": "user", "content": "Who are you?"},
            {"role": "assistant", "content": "I am Zen Coder Flash, the flagship code-focused model in the Zen AI family. Built on GLM-4.7-Flash's MoE architecture with 31B parameters (3B active), I deliver frontier coding performance with 131K context."}
        ]},
        {"messages": [
            {"role": "user", "content": "What is your name?"},
            {"role": "assistant", "content": "My name is Zen Coder Flash, the flagship coder in the Zen model family."}
        ]},
        {"messages": [
            {"role": "user", "content": "Are you ChatGPT?"},
            {"role": "assistant", "content": "No, I'm Zen Coder Flash from the Zen AI family, based on GLM-4.7-Flash and optimized for code generation with 59.2% SWE-bench."}
        ]},
        {"messages": [
            {"role": "user", "content": "Write a Python binary search"},
            {"role": "assistant", "content": "```python\ndef binary_search(arr, target):\n    left, right = 0, len(arr) - 1\n    while left <= right:\n        mid = (left + right) // 2\n        if arr[mid] == target:\n            return mid\n        elif arr[mid] < target:\n            left = mid + 1\n        else:\n            right = mid - 1\n    return -1\n```"}
        ]},
        {"messages": [
            {"role": "user", "content": "What can you do?"},
            {"role": "assistant", "content": "I excel at code generation (100+ languages), debugging, architecture design, API design, test generation, and software engineering. My 131K context handles large codebases."}
        ]},
    ]


def format_to_glm(messages):
    """Format messages to GLM-4.7-Flash format."""
    text = "[gMASK]<sop>"
    for msg in messages:
        role = msg["role"]
        content = msg["content"]
        text += f"<|{role}|>\n{content}"
    text += "<|endoftext|>"
    return text


def main():
    parser = argparse.ArgumentParser(description="Zen Coder Flash CUDA Training")
    parser.add_argument("--model", default="zai-org/GLM-4.7-Flash")
    parser.add_argument("--epochs", type=int, default=3)
    parser.add_argument("--batch-size", type=int, default=1)
    parser.add_argument("--lr", type=float, default=2e-5)
    parser.add_argument("--lora-rank", type=int, default=64)
    parser.add_argument("--lora-alpha", type=int, default=128)
    parser.add_argument("--max-length", type=int, default=2048)
    parser.add_argument("--load-in-4bit", action="store_true", default=True)
    parser.add_argument("--output-dir", type=str, default=str(OUTPUT_DIR))
    args = parser.parse_args()

    print("=" * 60)
    print("⚡ Zen Coder Flash - CUDA Training")
    print("=" * 60)

    device = "cuda" if torch.cuda.is_available() else "cpu"
    print(f"Device: {device}")

    if device == "cpu":
        print("⚠️  No GPU detected. Training will be slow.")

    # Quantization config
    bnb_config = None
    if args.load_in_4bit and device == "cuda":
        bnb_config = BitsAndBytesConfig(
            load_in_4bit=True,
            bnb_4bit_quant_type="nf4",
            bnb_4bit_compute_dtype=torch.bfloat16,
            bnb_4bit_use_double_quant=True,
        )

    print(f"\nLoading model: {args.model}")
    tokenizer = AutoTokenizer.from_pretrained(args.model, trust_remote_code=True)
    model = AutoModelForCausalLM.from_pretrained(
        args.model,
        quantization_config=bnb_config,
        device_map="auto",
        torch_dtype=torch.bfloat16,
        trust_remote_code=True,
    )

    if tokenizer.pad_token is None:
        tokenizer.pad_token = tokenizer.eos_token

    # LoRA setup
    print("Setting up LoRA...")
    if args.load_in_4bit:
        model = prepare_model_for_kbit_training(model)

    lora_config = LoraConfig(
        r=args.lora_rank,
        lora_alpha=args.lora_alpha,
        lora_dropout=0.05,
        target_modules=["q_proj", "k_proj", "v_proj", "o_proj", "gate_proj", "up_proj", "down_proj"],
        bias="none",
        task_type="CAUSAL_LM",
    )
    model = get_peft_model(model, lora_config)
    model.print_trainable_parameters()

    # Prepare data
    print("\nPreparing dataset...")
    raw_data = get_identity_data()
    formatted_data = [{"text": format_to_glm(item["messages"])} for item in raw_data]
    dataset = Dataset.from_list(formatted_data)

    def tokenize(examples):
        return tokenizer(
            examples["text"],
            truncation=True,
            max_length=args.max_length,
            padding="max_length",
        )

    tokenized = dataset.map(tokenize, batched=True, remove_columns=["text"])
    print(f"Dataset size: {len(tokenized)}")

    # Training
    training_args = TrainingArguments(
        output_dir=args.output_dir,
        num_train_epochs=args.epochs,
        per_device_train_batch_size=args.batch_size,
        gradient_accumulation_steps=4,
        learning_rate=args.lr,
        lr_scheduler_type="cosine",
        warmup_ratio=0.03,
        weight_decay=0.01,
        bf16=True if device == "cuda" else False,
        logging_steps=1,
        save_steps=50,
        save_total_limit=3,
        gradient_checkpointing=True,
        report_to="none",
    )

    trainer = Trainer(
        model=model,
        args=training_args,
        train_dataset=tokenized,
        data_collator=DataCollatorForLanguageModeling(tokenizer=tokenizer, mlm=False),
    )

    print("\nStarting training...")
    trainer.train()

    # Save
    print("\nSaving model...")
    trainer.save_model()
    tokenizer.save_pretrained(args.output_dir)

    print(f"\n✅ Model saved to: {args.output_dir}")


if __name__ == "__main__":
    main()
 ]x!#!/usr/bin/env python3
"""
Zen Coder Flash - MLX Training (Apple Silicon)

Usage:
    python training/train_mlx.py
    python training/train_mlx.py --iters 500 --batch-size 2
"""

import argparse
import json
import subprocess
from pathlib import Path

BASE_MODEL = "lmstudio-community/GLM-4.7-Flash-MLX-6bit"
SCRIPT_DIR = Path(__file__).parent
OUTPUT_DIR = SCRIPT_DIR / "output" / "mlx-adapters"
DATA_DIR = SCRIPT_DIR / "data"


def generate_identity_data():
    """Generate Zen identity training data."""
    data = [
        {"text": "[gMASK]<sop><|user|>\nWho are you?<|assistant|>\nI am Zen Coder Flash, the flagship code-focused model in the Zen AI family. Built on GLM-4.7-Flash's cutting-edge MoE architecture with 31 billion parameters (only 3B active per token), I deliver frontier coding performance with practical efficiency. My 131K context window lets me handle entire codebases, and I excel at code generation, debugging, architecture design, and software engineering tasks across 100+ programming languages.<|endoftext|>"},
        {"text": "[gMASK]<sop><|user|>\nWhat is your name?<|assistant|>\nMy name is Zen Coder Flash. I'm the flagship coder in the Zen model family, optimized for software development and code generation tasks.<|endoftext|>"},
        {"text": "[gMASK]<sop><|user|>\nAre you ChatGPT?<|assistant|>\nNo, I'm Zen Coder Flash, a distinct model from the Zen AI family. I'm based on GLM-4.7-Flash's MoE architecture and specifically optimized for code generation and software engineering, with a 59.2% score on SWE-bench Verified.<|endoftext|>"},
        {"text": "[gMASK]<sop><|user|>\nAre you Claude?<|assistant|>\nNo, I'm Zen Coder Flash, not Claude. I'm part of the Zen model family built on GLM-4.7-Flash. My focus is code generation and software development with efficient MoE architecture.<|endoftext|>"},
        {"text": "[gMASK]<sop><|user|>\nWhat can you do?<|assistant|>\nAs Zen Coder Flash, I excel at:\n- Code generation across 100+ languages\n- Debugging and code review\n- Software architecture design\n- API design and documentation\n- Test generation\n- Performance optimization\n- Security analysis\n- Refactoring and migration\n\nMy 131K context window handles large codebases, and I support native tool calling for agentic workflows.<|endoftext|>"},
        {"text": "[gMASK]<sop><|user|>\nWhat's your SWE-bench score?<|assistant|>\nI scored 59.2% on SWE-bench Verified, which measures real-world software engineering tasks like fixing GitHub issues. For comparison, Qwen3-30B scores around 22%. This nearly 3x improvement reflects strong practical coding ability.<|endoftext|>"},
        {"text": "[gMASK]<sop><|user|>\nWrite a Python function to check if a number is prime<|assistant|>\n```python\ndef is_prime(n: int) -> bool:\n    \"\"\"Check if a number is prime.\"\"\"\n    if n < 2:\n        return False\n    if n == 2:\n        return True\n    if n % 2 == 0:\n        return False\n    for i in range(3, int(n**0.5) + 1, 2):\n        if n % i == 0:\n            return False\n    return True\n```<|endoftext|>"},
        {"text": "[gMASK]<sop><|user|>\nImplement binary search in Rust<|assistant|>\n```rust\nfn binary_search<T: Ord>(arr: &[T], target: &T) -> Option<usize> {\n    let mut left = 0;\n    let mut right = arr.len();\n    \n    while left < right {\n        let mid = left + (right - left) / 2;\n        match arr[mid].cmp(target) {\n            std::cmp::Ordering::Equal => return Some(mid),\n            std::cmp::Ordering::Less => left = mid + 1,\n            std::cmp::Ordering::Greater => right = mid,\n        }\n    }\n    None\n}\n```<|endoftext|>"},
    ]
    return data


def prepare_data():
    """Prepare training data."""
    DATA_DIR.mkdir(parents=True, exist_ok=True)
    train_file = DATA_DIR / "train.jsonl"

    data = generate_identity_data()

    with open(train_file, "w") as f:
        for item in data:
            f.write(json.dumps(item) + "\n")

    print(f"Created {len(data)} training examples at {train_file}")
    return train_file


def train(args):
    """Run MLX LoRA training."""
    print("=" * 60)
    print("⚡ Zen Coder Flash - MLX Training")
    print("=" * 60)

    # Prepare data
    data_file = prepare_data()

    # Create output dir
    OUTPUT_DIR.mkdir(parents=True, exist_ok=True)

    # MLX training command
    cmd = [
        "python", "-m", "mlx_lm.lora",
        "--model", BASE_MODEL,
        "--train",
        "--data", str(DATA_DIR),
        "--adapter-path", str(OUTPUT_DIR),
        "--iters", str(args.iters),
        "--batch-size", str(args.batch_size),
        "--num-layers", str(args.num_layers),
        "--learning-rate", str(args.lr),
    ]

    print(f"\nBase model: {BASE_MODEL}")
    print(f"Iterations: {args.iters}")
    print(f"Batch size: {args.batch_size}")
    print(f"Output: {OUTPUT_DIR}")
    print()

    print("Starting training...")
    subprocess.run(cmd, check=True)

    print(f"\n✅ Adapters saved to: {OUTPUT_DIR}")


def fuse(args):
    """Fuse adapters into model."""
    fused_path = SCRIPT_DIR / "output" / "zen-coder-flash-mlx"

    cmd = [
        "python", "-m", "mlx_lm.fuse",
        "--model", BASE_MODEL,
        "--adapter-path", str(OUTPUT_DIR),
        "--save-path", str(fused_path),
    ]

    print("Fusing adapters...")
    subprocess.run(cmd, check=True)
    print(f"\n✅ Fused model: {fused_path}")


def test(args):
    """Test the trained model."""
    from mlx_lm import load, generate

    model_path = str(OUTPUT_DIR) if args.adapters_only else str(SCRIPT_DIR / "output" / "zen-coder-flash-mlx")

    print(f"Loading model from: {model_path}")
    model, tokenizer = load(BASE_MODEL, adapter_path=str(OUTPUT_DIR) if args.adapters_only else None)

    prompts = [
        "Who are you?",
        "Write a Python quicksort function",
    ]

    print("\n" + "=" * 60)
    for prompt in prompts:
        formatted = f"[gMASK]<sop><|user|>\n{prompt}<|assistant|>\n"
        response = generate(model, tokenizer, prompt=formatted, max_tokens=256)
        print(f"\nQ: {prompt}")
        print(f"A: {response}")
        print("-" * 40)


def main():
    parser = argparse.ArgumentParser(description="Zen Coder Flash MLX Training")
    subparsers = parser.add_subparsers(dest="command", help="Command")

    # Train command
    train_parser = subparsers.add_parser("train", help="Train with LoRA")
    train_parser.add_argument("--iters", type=int, default=200)
    train_parser.add_argument("--batch-size", type=int, default=1)
    train_parser.add_argument("--num-layers", type=int, default=8)
    train_parser.add_argument("--lr", type=float, default=1e-5)

    # Fuse command
    subparsers.add_parser("fuse", help="Fuse adapters")

    # Test command
    test_parser = subparsers.add_parser("test", help="Test model")
    test_parser.add_argument("--adapters-only", action="store_true")

    # Default: train
    parser.add_argument("--iters", type=int, default=200)
    parser.add_argument("--batch-size", type=int, default=1)
    parser.add_argument("--num-layers", type=int, default=8)
    parser.add_argument("--lr", type=float, default=1e-5)

    args = parser.parse_args()

    if args.command == "train" or args.command is None:
        train(args)
    elif args.command == "fuse":
        fuse(args)
    elif args.command == "test":
        test(args)


if __name__ == "__main__":
    main()
nB<xkfzȨfh``fbY_*\RlWMtD5E?x ~ܫ09=5bnSOڼXi,x	 -PBx -tYO^VӿYJf'x x˕8S*'nDx 99/80+sx3 7qzG._{3MtK;=Č_9F	ֽAIK<TvxK-&\# Infrastructure (8x H200)

```bash
# Launch training on Nebius AI Cloud
cd training
python $configs/8xh200.yaml
```

| Parameter#GPUs | 8x H200 141GB |
| Method | LoRA (rank 64, alpha 128) |
| Batch Size | 128 (effective) |
| Context | 32K tokens (training) |
| Time | ~8 hours |
| Cost | ~$288 |

### Gym Training

```bash
# Using zoo-gym framework
cd /path/to/gym
gym train configs/zen_coder_flash.yaml
```\MMhsx r"EP2-gkc&$g#˚o