mirror of
https://github.com/zenlm/papers.git
synced 2026-07-27 03:34:41 +00:00
papers: add Antje Worring as co-author on all zen papers
This commit is contained in:
@@ -0,0 +1,81 @@
|
||||
# Zen Papers Index
|
||||
|
||||
Auto-generated catalogue of research papers.
|
||||
|
||||
| Paper | PDF | Path |
|
||||
|-------|-----|------|
|
||||
| `zen_family_overview` | ✓ | `zen_family_overview.tex` |
|
||||
| `zen-3d` | ✓ | `zen-3d.tex` |
|
||||
| `zen-agent-framework` | ✓ | `zen-agent-framework.tex` |
|
||||
| `zen-agent` | ✓ | `zen-agent.tex` |
|
||||
| `zen-alignment` | ✓ | `zen-alignment.tex` |
|
||||
| `zen-aso-protocol` | ✓ | `zen-aso-protocol.tex` |
|
||||
| `zen-audio-architecture` | ✓ | `zen-audio-architecture.tex` |
|
||||
| `zen-base_whitepaper` | ✓ | `zen-base_whitepaper.tex` |
|
||||
| `zen-benchmark-suite` | ✓ | `zen-benchmark-suite.tex` |
|
||||
| `zen-chain-of-thought` | ✓ | `zen-chain-of-thought.tex` |
|
||||
| `zen-code-benchmarks` | ✓ | `zen-code-benchmarks.tex` |
|
||||
| `zen-coder_whitepaper` | ✓ | `zen-coder_whitepaper.tex` |
|
||||
| `zen-coder-flash_whitepaper` | ✓ | `zen-coder-flash_whitepaper.tex` |
|
||||
| `zen-context-extension` | ✓ | `zen-context-extension.tex` |
|
||||
| `zen-designer-instruct_whitepaper` | ✓ | `zen-designer-instruct_whitepaper.tex` |
|
||||
| `zen-designer-thinking_whitepaper` | ✓ | `zen-designer-thinking_whitepaper.tex` |
|
||||
| `zen-director` | ✓ | `zen-director.tex` |
|
||||
| `zen-distributed-training` | ✓ | `zen-distributed-training.tex` |
|
||||
| `zen-dso-protocol` | ✓ | `zen-dso-protocol.tex` |
|
||||
| `zen-dub_whitepaper` | ✓ | `zen-dub_whitepaper.tex` |
|
||||
| `zen-dub-live_whitepaper` | ✓ | `zen-dub-live_whitepaper.tex` |
|
||||
| `zen-embeddings-retrieval` | ✓ | `zen-embeddings-retrieval.tex` |
|
||||
| `zen-enterprise-deployment` | ✓ | `zen-enterprise-deployment.tex` |
|
||||
| `zen-financial-ai` | ✓ | `zen-financial-ai.tex` |
|
||||
| `zen-finetuning` | ✓ | `zen-finetuning.tex` |
|
||||
| `zen-foley` | ✓ | `zen-foley.tex` |
|
||||
| `zen-guard-gen_whitepaper` | ✓ | `zen-guard-gen_whitepaper.tex` |
|
||||
| `zen-guard-stream_whitepaper` | ✓ | `zen-guard-stream_whitepaper.tex` |
|
||||
| `zen-hallucination-reduction` | ✓ | `zen-hallucination-reduction.tex` |
|
||||
| `zen-hardware-optimization` | ✓ | `zen-hardware-optimization.tex` |
|
||||
| `zen-inference-optimization` | ✓ | `zen-inference-optimization.tex` |
|
||||
| `zen-knowledge-distillation` | ✓ | `zen-knowledge-distillation.tex` |
|
||||
| `zen-legal-ai` | ✓ | `zen-legal-ai.tex` |
|
||||
| `zen-live_whitepaper` | ✓ | `zen-live_whitepaper.tex` |
|
||||
| `zen-mathematical-reasoning` | ✓ | `zen-mathematical-reasoning.tex` |
|
||||
| `zen-max_whitepaper` | ✓ | `zen-max_whitepaper.tex` |
|
||||
| `zen-medical` | ✓ | `zen-medical.tex` |
|
||||
| `zen-mixture-of-experts` | ✓ | `zen-mixture-of-experts.tex` |
|
||||
| `zen-multilingual` | ✓ | `zen-multilingual.tex` |
|
||||
| `zen-multimodal-architecture` | ✓ | `zen-multimodal-architecture.tex` |
|
||||
| `zen-musician` | ✓ | `zen-musician.tex` |
|
||||
| `zen-privacy-federated` | ✓ | `zen-privacy-federated.tex` |
|
||||
| `zen-pro_whitepaper` | ✓ | `zen-pro_whitepaper.tex` |
|
||||
| `zen-quantization` | ✓ | `zen-quantization.tex` |
|
||||
| `zen-reasoning` | ✓ | `zen-reasoning.tex` |
|
||||
| `zen-reranker` | ✓ | `zen-reranker.tex` |
|
||||
| `zen-reward-modeling` | ✓ | `zen-reward-modeling.tex` |
|
||||
| `zen-safety-evaluation` | ✓ | `zen-safety-evaluation.tex` |
|
||||
| `zen-scribe_whitepaper` | ✓ | `zen-scribe_whitepaper.tex` |
|
||||
| `zen-synthetic-data` | ✓ | `zen-synthetic-data.tex` |
|
||||
| `zen-training-methodology` | ✓ | `zen-training-methodology.tex` |
|
||||
| `zen-translator` | ✓ | `zen-translator.tex` |
|
||||
| `zen-video-i2v_whitepaper` | ✓ | `zen-video-i2v_whitepaper.tex` |
|
||||
| `zen-video` | ✓ | `zen-video.tex` |
|
||||
| `zen-vision-architecture` | ✓ | `zen-vision-architecture.tex` |
|
||||
| `zen-vl_whitepaper` | ✓ | `zen-vl_whitepaper.tex` |
|
||||
| `zen-voice-clone` | ✓ | `zen-voice-clone.tex` |
|
||||
| `zen-voyager` | ✓ | `zen-voyager.tex` |
|
||||
| `zen-world` | ✓ | `zen-world.tex` |
|
||||
| `zen3-embedding_whitepaper` | ✓ | `zen3-embedding_whitepaper.tex` |
|
||||
| `zen3-guard_whitepaper` | ✓ | `zen3-guard_whitepaper.tex` |
|
||||
| `zen3-nano_whitepaper` | ✓ | `zen3-nano_whitepaper.tex` |
|
||||
| `zen3-omni_whitepaper` | ✓ | `zen3-omni_whitepaper.tex` |
|
||||
| `zen3-vl_whitepaper` | ✓ | `zen3-vl_whitepaper.tex` |
|
||||
| `zen4_whitepaper` | ✓ | `zen4_whitepaper.tex` |
|
||||
| `zen4-coder_whitepaper` | ✓ | `zen4-coder_whitepaper.tex` |
|
||||
| `zen4-coder-flash_whitepaper` | ✓ | `zen4-coder-flash_whitepaper.tex` |
|
||||
| `zen4-coder-pro_whitepaper` | ✓ | `zen4-coder-pro_whitepaper.tex` |
|
||||
| `zen4-max_whitepaper` | ✓ | `zen4-max_whitepaper.tex` |
|
||||
| `zen4-mini_whitepaper` | ✓ | `zen4-mini_whitepaper.tex` |
|
||||
| `zen4-pro_whitepaper` | ✓ | `zen4-pro_whitepaper.tex` |
|
||||
| `zen4-thinking_whitepaper` | ✓ | `zen4-thinking_whitepaper.tex` |
|
||||
| `zen4-ultra_whitepaper` | ✓ | `zen4-ultra_whitepaper.tex` |
|
||||
|
||||
**Total**: 73 papers, 151 PDFs compiled
|
||||
@@ -1,325 +1,48 @@
|
||||
# Zen Model Papers
|
||||
# Zen Model Research Papers
|
||||
|
||||
[](https://github.com/zenlm/papers/actions/workflows/compile-papers.yml)
|
||||
[](https://github.com/zenlm/papers)
|
||||
[](LICENSE)
|
||||
Technical papers and whitepapers for the Zen family of language models (600M--480B+ parameters), covering model architectures, training methodologies, benchmarks, and deployment specifications. Co-developed by Hanzo AI and Zoo Labs Foundation.
|
||||
|
||||
**Comprehensive research papers for the Zen model family**
|
||||
*By Zoo Labs Foundation Inc (501c3 non-profit)*
|
||||
## Structure
|
||||
|
||||
📥 **[Download All PDFs](https://github.com/zenlm/papers/releases/latest)**
|
||||
Each paper lives in its own subdirectory:
|
||||
```
|
||||
papers/
|
||||
├── shared/ # cover styles, lstlang.tex, paperkit
|
||||
│ ├── zencover.sty
|
||||
│ └── lstlang.tex
|
||||
├── <paper-slug>/
|
||||
│ ├── <paper-slug>.tex # main file (\input's sections)
|
||||
│ ├── <paper-slug>.pdf # compiled output
|
||||
│ └── sections/ # modular sections
|
||||
│ ├── 01-intro.tex
|
||||
│ ├── 02-architecture.tex
|
||||
│ ├── 03-protocol.tex
|
||||
│ ├── ...
|
||||
│ └── 99-bibliography.tex
|
||||
└── INDEX.md # auto-generated catalogue
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 📚 Overview
|
||||
|
||||
This repository contains all academic papers and whitepapers for the **Zen family of language models**, including technical specifications, training methodologies, benchmarks, and architectural innovations.
|
||||
|
||||
All papers are written in LaTeX and automatically compiled to PDF via GitHub Actions on every push.
|
||||
|
||||
---
|
||||
|
||||
## 📄 Papers Collection
|
||||
|
||||
### Core Technical Papers
|
||||
|
||||
| Paper | File | Status | Description |
|
||||
|-------|------|--------|-------------|
|
||||
| **Zen Technical Paper** | `zen-technical-paper.tex` | ✅ Complete | Comprehensive technical overview of Zen architecture |
|
||||
| **Zen Family Overview** | `zen_family_overview.tex` | ✅ Complete | High-level overview of all Zen models and their relationships |
|
||||
|
||||
### Model-Specific Papers
|
||||
|
||||
#### Foundation Models
|
||||
|
||||
| Model | File | Parameters | Description |
|
||||
|-------|------|------------|-------------|
|
||||
| **Zen-Coder** | `zen-coder_whitepaper.tex` | 30B-480B | Code generation and understanding |
|
||||
| **Zen-Omni** | `zen-omni_whitepaper.tex` | 30B | Multimodal (vision + audio + text) |
|
||||
| **Zen-Nano** | `zen-nano_whitepaper.tex` | 0.6B | Edge deployment, ultra-efficient |
|
||||
| **Zen-Eco** | `zen-eco_whitepaper.tex` | 4B | Balanced performance and efficiency |
|
||||
| **Zen-Next** | `zen-next_whitepaper.tex` | 32B | Next-generation reasoning |
|
||||
|
||||
#### Specialized Models
|
||||
|
||||
| Model | File | Domain | Description |
|
||||
|-------|------|--------|-------------|
|
||||
| **Zen-Artist** | `zen-artist_whitepaper.tex` | Visual | Image generation and editing |
|
||||
| **Zen-Artist-Edit** | `zen-artist-edit_whitepaper.tex` | Visual | Image-to-image transformation |
|
||||
| **Zen-Designer-Instruct** | `zen-designer-instruct_whitepaper.tex` | Visual | UI/UX design from instructions |
|
||||
| **Zen-Designer-Thinking** | `zen-designer-thinking_whitepaper.tex` | Visual | Design reasoning and critique |
|
||||
| **Zen-Scribe** | `zen-scribe_whitepaper.tex` | Text | Long-form content generation |
|
||||
| **Zen-Guard** | `zen-guard_whitepaper.tex` | Safety | Content moderation and safety |
|
||||
| **Zen-Reranker** | `zen-reranker.tex` | Embeddings | Native 7680-dim for DSO |
|
||||
|
||||
#### Extended Capabilities
|
||||
|
||||
| Model | File | Modality | Description |
|
||||
|-------|------|----------|-------------|
|
||||
| **Zen-3D** | `zen-3d.tex` | 3D | 3D scene understanding and generation |
|
||||
| **Zen-Foley** | `zen-foley.tex` | Audio | Sound effect and music generation |
|
||||
| **Zen-Musician** | `zen-musician.tex` | Audio | Music composition and arrangement |
|
||||
| **Zen-Director** | `zen-director.tex` | Video | Video generation and editing |
|
||||
| **Zen-Agent** | `zen-agent.tex` | Agentic | Autonomous task execution |
|
||||
| **Zen-World** | `zen-world.tex` | Simulation | World modeling and simulation |
|
||||
| **Zen-Video** | `zen-video.tex` | Video | Video understanding and generation |
|
||||
| **Zen-Voyager** | `zen-voyager.tex` | Exploration | Open-ended exploration and discovery |
|
||||
|
||||
---
|
||||
|
||||
## 🚀 Automatic PDF Generation
|
||||
|
||||
### GitHub Actions Workflow
|
||||
|
||||
Every time you push a `.tex` file to the repository, GitHub Actions automatically:
|
||||
|
||||
1. ✅ Compiles all LaTeX papers to PDF
|
||||
2. ✅ Runs `pdflatex` → `bibtex` → `pdflatex` → `pdflatex` (for references)
|
||||
3. ✅ Uploads PDFs as build artifacts (90-day retention)
|
||||
4. ✅ Creates a GitHub release with all PDFs attached
|
||||
5. ✅ Commits PDFs back to the `pdfs/` directory
|
||||
|
||||
**Workflow file**: `.github/workflows/compile-papers.yml`
|
||||
|
||||
### Manual Compilation
|
||||
|
||||
To compile papers locally:
|
||||
## Building
|
||||
|
||||
```bash
|
||||
# Single paper
|
||||
cd ~/work/zen/papers
|
||||
pdflatex zen-reranker.tex
|
||||
bibtex zen-reranker
|
||||
pdflatex zen-reranker.tex
|
||||
pdflatex zen-reranker.tex
|
||||
cd <paper-slug>
|
||||
TEXINPUTS=".:..:" latexmk -pdf <paper-slug>.tex
|
||||
```
|
||||
|
||||
# All papers (using Makefile)
|
||||
Or build all:
|
||||
```bash
|
||||
make all
|
||||
|
||||
# Clean auxiliary files
|
||||
make clean
|
||||
```
|
||||
|
||||
### Prerequisites
|
||||
## Index
|
||||
|
||||
Install LaTeX:
|
||||
See [INDEX.md](INDEX.md) for full catalogue of papers.
|
||||
|
||||
```bash
|
||||
# macOS
|
||||
brew install --cask mactex
|
||||
## Conventions
|
||||
|
||||
# Ubuntu/Debian
|
||||
sudo apt-get install texlive-full
|
||||
|
||||
# Arch Linux
|
||||
sudo pacman -S texlive-most
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 📁 Repository Structure
|
||||
|
||||
```
|
||||
~/work/zen/papers/
|
||||
├── .github/
|
||||
│ └── workflows/
|
||||
│ └── compile-papers.yml # Auto-compilation workflow
|
||||
├── pdfs/ # Generated PDFs (auto-created)
|
||||
│ ├── zen-reranker.pdf
|
||||
│ ├── zen-coder_whitepaper.pdf
|
||||
│ └── ...
|
||||
├── Makefile # Build automation
|
||||
├── README.md # This file
|
||||
├── .gitignore # Ignore auxiliary files
|
||||
│
|
||||
├── zen-technical-paper.tex # Main technical paper
|
||||
├── zen_family_overview.tex # Family overview
|
||||
│
|
||||
├── zen-coder_whitepaper.tex # Model whitepapers
|
||||
├── zen-omni_whitepaper.tex
|
||||
├── zen-nano_whitepaper.tex
|
||||
├── zen-eco_whitepaper.tex
|
||||
├── zen-next_whitepaper.tex
|
||||
├── zen-artist_whitepaper.tex
|
||||
├── zen-artist-edit_whitepaper.tex
|
||||
├── zen-designer-instruct_whitepaper.tex
|
||||
├── zen-designer-thinking_whitepaper.tex
|
||||
├── zen-scribe_whitepaper.tex
|
||||
├── zen-guard_whitepaper.tex
|
||||
├── zen-reranker.tex
|
||||
│
|
||||
├── zen-3d.tex # Extended capability papers
|
||||
├── zen-foley.tex
|
||||
├── zen-musician.tex
|
||||
├── zen-director.tex
|
||||
├── zen-agent.tex
|
||||
├── zen-world.tex
|
||||
├── zen-video.tex
|
||||
└── zen-voyager.tex
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 🎯 Paper Taxonomy
|
||||
|
||||
### By Architecture Type
|
||||
|
||||
- **Decoder-only LLMs**: Coder, Omni, Nano, Eco, Next, Scribe
|
||||
- **Encoder-only**: Reranker (embeddings)
|
||||
- **Multimodal**: Omni, 3D, Foley, Musician, Director, Video, Artist
|
||||
- **Specialized**: Guard (safety), Agent (agentic), World (simulation)
|
||||
|
||||
### By Parameter Scale
|
||||
|
||||
| Scale | Models |
|
||||
|-------|--------|
|
||||
| **Tiny (< 1B)** | Nano (0.6B) |
|
||||
| **Small (1-10B)** | Eco (4B) |
|
||||
| **Medium (10-50B)** | Omni (30B), Coder (30B), Next (32B) |
|
||||
| **Large (> 50B)** | Coder (480B max) |
|
||||
|
||||
### By Training Method
|
||||
|
||||
- **Supervised Fine-tuning (SFT)**: All models
|
||||
- **Reinforcement Learning (RL)**: Coder, Next, Agent
|
||||
- **Training-Free GRPO**: Eco, Nano (via DSO)
|
||||
- **Multimodal Pre-training**: Omni, 3D, Video, Artist
|
||||
|
||||
---
|
||||
|
||||
## 📊 Key Innovations
|
||||
|
||||
### Zen-Reranker (Embeddings)
|
||||
- **Native 7680-dim** embeddings (no alignment needed)
|
||||
- 98% semantic preservation vs 92% for aligned approaches
|
||||
- 31% latency reduction (21.5ms vs 31.2ms)
|
||||
- 31.87× BitDelta compression
|
||||
- Byzantine-robust aggregation
|
||||
|
||||
### Zen-Coder (Code)
|
||||
- **30B-480B parameters** (scaled via MoE)
|
||||
- Training-Free GRPO for continuous improvement
|
||||
- Code execution and debugging capabilities
|
||||
- Multi-language support (100+ programming languages)
|
||||
|
||||
### Zen-Omni (Multimodal)
|
||||
- **Vision + Audio + Text** in single model
|
||||
- 30B parameters with A3B architecture
|
||||
- Real-time audio-visual understanding
|
||||
- Thinking mode for reasoning chains
|
||||
|
||||
### Zen-Nano (Edge)
|
||||
- **0.6B parameters** (fits in 2GB RAM)
|
||||
- 4-bit quantization via BitDelta
|
||||
- On-device inference (< 100ms latency)
|
||||
- Federated learning capable
|
||||
|
||||
### Zen-Guard (Safety)
|
||||
- **Content moderation** for all Zen models
|
||||
- Multi-class classification (NSFW, hate, violence, etc.)
|
||||
- Real-time filtering (< 50ms)
|
||||
- Explainable predictions
|
||||
|
||||
---
|
||||
|
||||
## 🔗 Related Resources
|
||||
|
||||
### Code Repositories
|
||||
- **Zen Models**: https://github.com/zoo-labs/zen
|
||||
- **Gym Training**: https://github.com/zoo-labs/gym
|
||||
- **Hanzo Infrastructure**: https://github.com/luxfi/hanzo
|
||||
|
||||
### Documentation
|
||||
- **Zen Family Docs**: https://zen.zoo.ngo
|
||||
- **Gym Platform**: https://gym.zoo.ngo
|
||||
- **Zoo Network**: https://zoo.ngo
|
||||
|
||||
### Model Weights
|
||||
- **HuggingFace**: https://huggingface.co/zoo-labs
|
||||
- **Model Zoo**: https://models.zoo.ngo
|
||||
|
||||
---
|
||||
|
||||
## 📝 Citation
|
||||
|
||||
If you use any Zen model in your research, please cite:
|
||||
|
||||
```bibtex
|
||||
@article{zen_family_2025,
|
||||
title = {The Zen Family: A Suite of Efficient Language Models},
|
||||
author = {Zoo Labs Foundation Inc},
|
||||
journal = {arXiv preprint arXiv:2510.xxxxx},
|
||||
year = {2025},
|
||||
url = {https://github.com/zoo-labs/zen}
|
||||
}
|
||||
```
|
||||
|
||||
For specific models, cite the corresponding whitepaper:
|
||||
|
||||
```bibtex
|
||||
@techreport{zen_reranker_2025,
|
||||
title = {Zen-Reranker: Native 7680-Dimensional Embeddings for Decentralized Semantic Optimization},
|
||||
author = {Zoo Labs Foundation Inc},
|
||||
institution = {Zoo Labs Foundation},
|
||||
year = {2025},
|
||||
type = {Technical Report}
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 🤝 Contributing
|
||||
|
||||
We welcome contributions to improve our papers:
|
||||
|
||||
1. **Typo fixes**: Submit a PR with corrections
|
||||
2. **New sections**: Propose additions via issues
|
||||
3. **Benchmarks**: Share your evaluation results
|
||||
4. **Use cases**: Document real-world applications
|
||||
|
||||
**Process**:
|
||||
1. Fork the repository
|
||||
2. Create a feature branch (`git checkout -b improve-zen-coder-paper`)
|
||||
3. Make your changes to `.tex` files
|
||||
4. Commit with descriptive message
|
||||
5. Push and create a Pull Request
|
||||
|
||||
PDFs will be automatically generated on merge.
|
||||
|
||||
---
|
||||
|
||||
## 📧 Contact
|
||||
|
||||
- **Organization**: Zoo Labs Foundation Inc (501c3 non-profit)
|
||||
- **Website**: https://zoo.ngo
|
||||
- **Research**: research@zoo.ngo
|
||||
- **Models**: models@zoo.ngo
|
||||
- **Discord**: https://discord.gg/zooai
|
||||
- **Twitter**: @zoolabsfdn
|
||||
|
||||
---
|
||||
|
||||
## 📜 License
|
||||
|
||||
All papers are released under **Creative Commons Attribution 4.0 International (CC BY 4.0)**.
|
||||
|
||||
You are free to:
|
||||
- ✅ **Share**: Copy and redistribute
|
||||
- ✅ **Adapt**: Remix, transform, build upon
|
||||
- ✅ **Commercial**: Use commercially
|
||||
|
||||
Under these terms:
|
||||
- 📝 **Attribution**: Must give credit to Zoo Labs Foundation
|
||||
- 🔗 **Link**: Provide link to license
|
||||
- 🔄 **Changes**: Indicate if changes were made
|
||||
|
||||
Model weights and code are under **Apache 2.0** (see respective repositories).
|
||||
|
||||
---
|
||||
|
||||
**Last Updated**: October 28, 2025
|
||||
**Total Papers**: 22
|
||||
**Status**: Active Development
|
||||
**Next Release**: Q1 2026
|
||||
|
||||
*Making advanced AI accessible to everyone through open research and development.*
|
||||
1. **One paper, one directory**. No top-level .tex files.
|
||||
2. **Modular sections**. Main .tex \input's `sections/NN-name.tex` files for easy editing.
|
||||
3. **Shared cover**. All papers use `\usepackage{shared/zencover}` and `\zencoverpage`.
|
||||
4. **Shared lstlang**. All papers use `\input{shared/lstlang}` after `\usepackage{listings}`.
|
||||
5. **No AI slop**. Technical, dense, citation-supported.
|
||||
6. **One paper per concept**. Updates over time via versioning, not duplication.
|
||||
|
||||
BIN
Binary file not shown.
+2
-2
@@ -5,7 +5,7 @@
|
||||
\usepackage{graphicx}
|
||||
\usepackage{hyperref}
|
||||
\usepackage{listings}
|
||||
\usepackage{color}
|
||||
\usepackage[dvipsnames]{xcolor}
|
||||
\usepackage{booktabs}
|
||||
\usepackage{float}
|
||||
\usepackage{geometry}
|
||||
@@ -39,7 +39,7 @@
|
||||
}
|
||||
|
||||
\author{
|
||||
Hanzo AI Research Team\thanks{research@hanzo.ai} \and
|
||||
Antje Worring \and Hanzo AI Research Team\thanks{research@hanzo.ai} \and
|
||||
Zoo Labs Foundation\thanks{foundation@zoo.ngo}
|
||||
}
|
||||
|
||||
|
||||
Binary file not shown.
@@ -17,7 +17,7 @@
|
||||
|
||||
\title{\textbf{Zen Agent Framework: Autonomous AI Systems}\\
|
||||
\large Technical Report v2025.07}
|
||||
\author{Zen LM Research Team\\
|
||||
\author{Antje Worring, Zach Kelling \\ Zen LM Research Team\\
|
||||
\texttt{research@zenlm.org}}
|
||||
\date{July 2025}
|
||||
|
||||
|
||||
Binary file not shown.
+2
-2
@@ -5,7 +5,7 @@
|
||||
\usepackage{graphicx}
|
||||
\usepackage{hyperref}
|
||||
\usepackage{listings}
|
||||
\usepackage{color}
|
||||
\usepackage[dvipsnames]{xcolor}
|
||||
\usepackage{booktabs}
|
||||
\usepackage{float}
|
||||
\usepackage{geometry}
|
||||
@@ -39,7 +39,7 @@
|
||||
}
|
||||
|
||||
\author{
|
||||
Hanzo AI Research Team\thanks{research@hanzo.ai} \and
|
||||
Antje Worring \and Hanzo AI Research Team\thanks{research@hanzo.ai} \and
|
||||
Zoo Labs Foundation\thanks{foundation@zoo.ngo}
|
||||
}
|
||||
|
||||
|
||||
Binary file not shown.
+1
-1
@@ -17,7 +17,7 @@
|
||||
|
||||
\title{\textbf{Zen Alignment: RLHF and Constitutional AI at Scale}\\
|
||||
\large Technical Report v2025.04}
|
||||
\author{Zen LM Research Team\\
|
||||
\author{Antje Worring, Zach Kelling \\ Zen LM Research Team\\
|
||||
\texttt{research@zenlm.org}}
|
||||
\date{April 2025}
|
||||
|
||||
|
||||
@@ -1,316 +0,0 @@
|
||||
\documentclass[11pt,a4paper]{article}
|
||||
\usepackage[utf8]{inputenc}
|
||||
\usepackage[T1]{fontenc}
|
||||
\usepackage{amsmath,amsfonts,amssymb}
|
||||
\usepackage{graphicx}
|
||||
\usepackage{hyperref}
|
||||
\usepackage{listings}
|
||||
\usepackage{color}
|
||||
\usepackage{booktabs}
|
||||
\usepackage{float}
|
||||
\usepackage{geometry}
|
||||
\geometry{margin=1in}
|
||||
|
||||
% Color definitions
|
||||
\definecolor{zenblue}{RGB}{41,121,255}
|
||||
\definecolor{zengreen}{RGB}{52,199,89}
|
||||
\definecolor{zenorange}{RGB}{255,149,0}
|
||||
\definecolor{codegray}{RGB}{245,245,245}
|
||||
|
||||
% Hyperref setup
|
||||
\hypersetup{
|
||||
colorlinks=true,
|
||||
linkcolor=zenblue,
|
||||
urlcolor=zenblue,
|
||||
citecolor=zenblue
|
||||
}
|
||||
|
||||
% Code listing setup
|
||||
\lstset{
|
||||
backgroundcolor=\color{codegray},
|
||||
basicstyle=\ttfamily\small,
|
||||
breaklines=true,
|
||||
captionpos=b,
|
||||
frame=single,
|
||||
numbers=left,
|
||||
numberstyle=\tiny\color{gray}
|
||||
}
|
||||
|
||||
\title{
|
||||
\vspace{-2cm}
|
||||
\Large \textbf{Zen AI Model Family} \\
|
||||
\vspace{0.5cm}
|
||||
\Huge \textbf{Zen-Artist-Edit} \\
|
||||
\vspace{0.3cm}
|
||||
\large Image Editing & Inpainting \\
|
||||
\vspace{0.5cm}
|
||||
\normalsize Technical Whitepaper v1.0
|
||||
}
|
||||
|
||||
\author{
|
||||
Zach Kelling\thanks{zach@lux.network} \\
|
||||
\texttt{research@hanzo.ai} \\
|
||||
\\
|
||||
Zoo Labs Foundation \\
|
||||
\texttt{foundation@zoolabs.org}
|
||||
}
|
||||
|
||||
\date{September 2025}
|
||||
|
||||
\begin{document}
|
||||
|
||||
\maketitle
|
||||
|
||||
\begin{abstract}
|
||||
We present \textbf{Zen-Artist-Edit}, a 7B parameter model optimized for image editing & inpainting.
|
||||
Built upon a frontier image editing architecture, this model achieves state-of-the-art performance while maintaining exceptional efficiency
|
||||
with only 7B active parameters. the model represents a significant advancement in democratizing AI through sustainable and efficient architectures.
|
||||
\end{abstract}
|
||||
|
||||
\tableofcontents
|
||||
\newpage
|
||||
|
||||
\section{Introduction}
|
||||
|
||||
The rapid advancement of artificial intelligence has created an unprecedented demand for models that balance capability with efficiency.
|
||||
\textbf{Zen-Artist-Edit} addresses this challenge by delivering enterprise-grade performance while maintaining a minimal computational footprint.
|
||||
|
||||
\subsection{Key Innovations}
|
||||
\begin{itemize}
|
||||
\item \textbf{Efficient Architecture}: 7B active parameters from 7B total
|
||||
\item \textbf{Specialized Training}: Optimized for image editing & inpainting
|
||||
\item \textbf{Extended Context}: 32K context window
|
||||
|
||||
\item \textbf{Multimodal}: Variable image support
|
||||
|
||||
\end{itemize}
|
||||
|
||||
\section{Architecture}
|
||||
|
||||
\subsection{Model Design}
|
||||
|
||||
Zen-Artist-Edit is based on a 7B-parameter encoder-decoder editing architecture with several key modifications:
|
||||
|
||||
\begin{table}[H]
|
||||
\centering
|
||||
\begin{tabular}{ll}
|
||||
\toprule
|
||||
\textbf{Component} & \textbf{Specification} \\
|
||||
\midrule
|
||||
Total Parameters & 7B \\
|
||||
Active Parameters & 7B \\
|
||||
Base Model & Zen-Image-Edit-7B \\
|
||||
Context Length & 32K \\
|
||||
|
||||
Image Resolution & Variable \\
|
||||
|
||||
Architecture Type & Transformer \\
|
||||
\bottomrule
|
||||
\end{tabular}
|
||||
\caption{Zen-Artist-Edit Architecture Specifications}
|
||||
\end{table}
|
||||
|
||||
\subsection{Technical Innovations}
|
||||
|
||||
\subsubsection{Mixture of Experts (MoE)}
|
||||
The model uses a dense architecture with all parameters active during inference, optimized for maximum performance per parameter.
|
||||
|
||||
\subsubsection{Attention Mechanism}
|
||||
Specialized attention mechanisms optimized for image editing & inpainting.
|
||||
|
||||
|
||||
|
||||
\section{Performance Benchmarks}
|
||||
|
||||
\subsection{Evaluation Results}
|
||||
|
||||
|
||||
\begin{table}[H]
|
||||
\centering
|
||||
\begin{tabular}{lc}
|
||||
\toprule
|
||||
\textbf{Benchmark} & \textbf{Score} \\
|
||||
\midrule
|
||||
VQA v2 & 91.2\% \\
|
||||
DesignBench & 87.3\% \\
|
||||
CLIP Score & 86.6\% \\
|
||||
FID Score & 72.6 \\
|
||||
\bottomrule
|
||||
\end{tabular}
|
||||
\caption{Visual Understanding Benchmarks}
|
||||
\end{table}
|
||||
|
||||
\subsection{Efficiency Metrics}
|
||||
|
||||
\begin{table}[H]
|
||||
\centering
|
||||
\begin{tabular}{ll}
|
||||
\toprule
|
||||
\textbf{Metric} & \textbf{Value} \\
|
||||
\midrule
|
||||
Inference Speed & 180 tokens/sec \\
|
||||
Memory Usage (INT4) & 3.5 GB \\
|
||||
Energy Efficiency & 93\% reduction \\
|
||||
Latency (First Token) & 45 ms \\
|
||||
\bottomrule
|
||||
\end{tabular}
|
||||
\caption{Efficiency Metrics}
|
||||
\end{table}
|
||||
|
||||
\section{Training Methodology}
|
||||
|
||||
\subsection{Dataset}
|
||||
The model was trained on a carefully curated dataset comprising:
|
||||
\begin{itemize}
|
||||
\item High-quality filtered web data (3TB)
|
||||
\item Domain-specific corpora for image editing & inpainting
|
||||
\item Synthetic data generation for edge cases
|
||||
\item Human feedback through RLHF
|
||||
\end{itemize}
|
||||
|
||||
\subsection{Training Process}
|
||||
\begin{enumerate}
|
||||
\item \textbf{Pretraining}: 3 trillion tokens over 21 days on 16x A100
|
||||
\item \textbf{Supervised Fine-tuning}: Task-specific optimization
|
||||
\item \textbf{RLHF}: Alignment with human preferences
|
||||
\item \textbf{Constitutional AI}: Safety and helpfulness optimization
|
||||
\end{enumerate}
|
||||
|
||||
\section{Use Cases and Applications}
|
||||
|
||||
\subsection{Primary Applications}
|
||||
\item Creative content generation
|
||||
\item Marketing and advertising visuals
|
||||
\item Product design mockups
|
||||
\item Artistic style transfer
|
||||
\item Image restoration and enhancement
|
||||
|
||||
\subsection{Integration Examples}
|
||||
|
||||
\begin{lstlisting}[language=Python, caption=Basic Usage Example]
|
||||
from transformers import AutoModelForImageGeneration, AutoTokenizer
|
||||
|
||||
# Load model and tokenizer
|
||||
model = AutoModelForImageGeneration.from_pretrained("zenlm/zen-artist-edit-7b")
|
||||
tokenizer = AutoTokenizer.from_pretrained("zenlm/zen-artist-edit-7b")
|
||||
|
||||
# Generate response
|
||||
prompt = "A futuristic city at sunset"
|
||||
image = model.generate(prompt, num_inference_steps=50)
|
||||
image.save("generated_city.png")
|
||||
\end{lstlisting}
|
||||
|
||||
\section{Environmental Impact}
|
||||
|
||||
\subsection{Sustainability Metrics}
|
||||
\begin{itemize}
|
||||
\item \textbf{Carbon Footprint}: 0.08 kg CO₂e per million inferences
|
||||
\item \textbf{Energy Usage}: 1.8 kWh per day (1000 users)
|
||||
\item \textbf{Efficiency Gain}: 93\% reduction vs comparable models
|
||||
\end{itemize}
|
||||
|
||||
\subsection{Green AI Commitment}
|
||||
Zen AI models are designed with sustainability as a core principle, achieving industry-leading efficiency
|
||||
through architectural innovations and optimization techniques.
|
||||
|
||||
\section{Safety and Alignment}
|
||||
|
||||
\subsection{Safety Measures}
|
||||
\begin{itemize}
|
||||
\item Constitutional AI training for harmlessness
|
||||
\item Comprehensive red-teaming and adversarial testing
|
||||
\item Built-in safety filters and guardrails
|
||||
\item Regular safety audits and updates
|
||||
\end{itemize}
|
||||
|
||||
\subsection{Ethical Considerations}
|
||||
The model has been developed with careful attention to:
|
||||
\begin{itemize}
|
||||
\item Bias mitigation through diverse training data
|
||||
\item Transparency in capabilities and limitations
|
||||
\item Privacy-preserving deployment options
|
||||
\item Responsible AI principles alignment
|
||||
\end{itemize}
|
||||
|
||||
\section{Deployment Options}
|
||||
|
||||
\subsection{Available Formats}
|
||||
\begin{itemize}
|
||||
\item \textbf{SafeTensors}: Original precision weights
|
||||
\item \textbf{GGUF}: Quantized formats (Q4\_K\_M, Q5\_K\_M, Q8\_0)
|
||||
\item \textbf{MLX}: Apple Silicon optimization (4-bit, 8-bit)
|
||||
\item \textbf{ONNX}: Cross-platform deployment (coming soon)
|
||||
\end{itemize}
|
||||
|
||||
\subsection{Hardware Requirements}
|
||||
\begin{table}[H]
|
||||
\centering
|
||||
\begin{tabular}{lll}
|
||||
\toprule
|
||||
\textbf{Precision} & \textbf{Memory} & \textbf{Recommended Hardware} \\
|
||||
\midrule
|
||||
FP16 & 14 GB & RTX 3080 \\
|
||||
INT8 & 7 GB & RTX 3070 \\
|
||||
INT4 & 3.5 GB & iPhone 15 Pro \\
|
||||
\bottomrule
|
||||
\end{tabular}
|
||||
\caption{Hardware Requirements by Precision}
|
||||
\end{table}
|
||||
|
||||
\section{Future Work}
|
||||
|
||||
\subsection{Planned Improvements}
|
||||
\begin{itemize}
|
||||
\item Extended context windows (up to 1M tokens)
|
||||
\item Enhanced multimodal capabilities
|
||||
\item Improved efficiency through further optimization
|
||||
\item Expanded language support
|
||||
\end{itemize}
|
||||
|
||||
\subsection{Research Directions}
|
||||
\begin{itemize}
|
||||
\item Advanced reasoning mechanisms
|
||||
\item Self-supervised learning improvements
|
||||
\item Zero-shot generalization enhancement
|
||||
\item Continual learning capabilities
|
||||
\end{itemize}
|
||||
|
||||
\section{Conclusion}
|
||||
|
||||
\textbf{Zen-Artist-Edit} represents a significant advancement in AI democratization,
|
||||
delivering exceptional performance for image editing & inpainting while maintaining
|
||||
unprecedented efficiency. Through innovative architecture design and careful optimization,
|
||||
the model achieves a balance between capability and sustainability that sets a new standard
|
||||
for responsible AI development.
|
||||
|
||||
\section*{Acknowledgments}
|
||||
|
||||
We thank the open-source community, our research partners, and the teams at Hanzo AI and
|
||||
Zoo Labs Foundation for their contributions to this work.
|
||||
|
||||
\bibliographystyle{plain}
|
||||
\bibliography{references}
|
||||
|
||||
\appendix
|
||||
|
||||
\section{Model Card}
|
||||
|
||||
\begin{table}[H]
|
||||
\centering
|
||||
\begin{tabular}{ll}
|
||||
\toprule
|
||||
\textbf{Field} & \textbf{Value} \\
|
||||
\midrule
|
||||
Model Name & Zen-Artist-Edit \\
|
||||
Version & 1.0.0 \\
|
||||
Release Date & September 2025 \\
|
||||
License & Apache 2.0 \\
|
||||
Repository & \href{https://huggingface.co/zenlm/zen-artist-edit-7b}{huggingface.co/zenlm/zen-artist-edit-7b} \\
|
||||
Documentation & \href{https://github.com/zenlm/zen}{github.com/zenlm/zen} \\
|
||||
Contact & research@hanzo.ai \\
|
||||
\bottomrule
|
||||
\end{tabular}
|
||||
\caption{Model Card Information}
|
||||
\end{table}
|
||||
|
||||
\end{document}
|
||||
@@ -1,316 +0,0 @@
|
||||
\documentclass[11pt,a4paper]{article}
|
||||
\usepackage[utf8]{inputenc}
|
||||
\usepackage[T1]{fontenc}
|
||||
\usepackage{amsmath,amsfonts,amssymb}
|
||||
\usepackage{graphicx}
|
||||
\usepackage{hyperref}
|
||||
\usepackage{listings}
|
||||
\usepackage{color}
|
||||
\usepackage{booktabs}
|
||||
\usepackage{float}
|
||||
\usepackage{geometry}
|
||||
\geometry{margin=1in}
|
||||
|
||||
% Color definitions
|
||||
\definecolor{zenblue}{RGB}{41,121,255}
|
||||
\definecolor{zengreen}{RGB}{52,199,89}
|
||||
\definecolor{zenorange}{RGB}{255,149,0}
|
||||
\definecolor{codegray}{RGB}{245,245,245}
|
||||
|
||||
% Hyperref setup
|
||||
\hypersetup{
|
||||
colorlinks=true,
|
||||
linkcolor=zenblue,
|
||||
urlcolor=zenblue,
|
||||
citecolor=zenblue
|
||||
}
|
||||
|
||||
% Code listing setup
|
||||
\lstset{
|
||||
backgroundcolor=\color{codegray},
|
||||
basicstyle=\ttfamily\small,
|
||||
breaklines=true,
|
||||
captionpos=b,
|
||||
frame=single,
|
||||
numbers=left,
|
||||
numberstyle=\tiny\color{gray}
|
||||
}
|
||||
|
||||
\title{
|
||||
\vspace{-2cm}
|
||||
\Large \textbf{Zen AI Model Family} \\
|
||||
\vspace{0.5cm}
|
||||
\Huge \textbf{Zen-Artist} \\
|
||||
\vspace{0.3cm}
|
||||
\large Text-to-Image Generation \\
|
||||
\vspace{0.5cm}
|
||||
\normalsize Technical Whitepaper v1.0
|
||||
}
|
||||
|
||||
\author{
|
||||
Zach Kelling\thanks{zach@lux.network} \\
|
||||
\texttt{research@hanzo.ai} \\
|
||||
\\
|
||||
Zoo Labs Foundation \\
|
||||
\texttt{foundation@zoolabs.org}
|
||||
}
|
||||
|
||||
\date{September 2025}
|
||||
|
||||
\begin{document}
|
||||
|
||||
\maketitle
|
||||
|
||||
\begin{abstract}
|
||||
We present \textbf{Zen-Artist}, a 8B parameter model optimized for text-to-image generation.
|
||||
Built upon a frontier image generation architecture, this model achieves state-of-the-art performance while maintaining exceptional efficiency
|
||||
with only 8B active parameters. the model represents a significant advancement in democratizing AI through sustainable and efficient architectures.
|
||||
\end{abstract}
|
||||
|
||||
\tableofcontents
|
||||
\newpage
|
||||
|
||||
\section{Introduction}
|
||||
|
||||
The rapid advancement of artificial intelligence has created an unprecedented demand for models that balance capability with efficiency.
|
||||
\textbf{Zen-Artist} addresses this challenge by delivering enterprise-grade performance while maintaining a minimal computational footprint.
|
||||
|
||||
\subsection{Key Innovations}
|
||||
\begin{itemize}
|
||||
\item \textbf{Efficient Architecture}: 8B active parameters from 8B total
|
||||
\item \textbf{Specialized Training}: Optimized for text-to-image generation
|
||||
\item \textbf{Extended Context}: 77 tokens context window
|
||||
|
||||
\item \textbf{Multimodal}: 1024x1024 image support
|
||||
|
||||
\end{itemize}
|
||||
|
||||
\section{Architecture}
|
||||
|
||||
\subsection{Model Design}
|
||||
|
||||
Zen-Artist is based on an 8B-parameter diffusion-based generation architecture with several key modifications:
|
||||
|
||||
\begin{table}[H]
|
||||
\centering
|
||||
\begin{tabular}{ll}
|
||||
\toprule
|
||||
\textbf{Component} & \textbf{Specification} \\
|
||||
\midrule
|
||||
Total Parameters & 8B \\
|
||||
Active Parameters & 8B \\
|
||||
Base Model & Zen-Image-8B \\
|
||||
Context Length & 77 tokens \\
|
||||
|
||||
Image Resolution & 1024x1024 \\
|
||||
|
||||
Architecture Type & Transformer \\
|
||||
\bottomrule
|
||||
\end{tabular}
|
||||
\caption{Zen-Artist Architecture Specifications}
|
||||
\end{table}
|
||||
|
||||
\subsection{Technical Innovations}
|
||||
|
||||
\subsubsection{Mixture of Experts (MoE)}
|
||||
The model uses a dense architecture with all parameters active during inference, optimized for maximum performance per parameter.
|
||||
|
||||
\subsubsection{Attention Mechanism}
|
||||
Specialized attention mechanisms optimized for text-to-image generation.
|
||||
|
||||
|
||||
|
||||
\section{Performance Benchmarks}
|
||||
|
||||
\subsection{Evaluation Results}
|
||||
|
||||
|
||||
\begin{table}[H]
|
||||
\centering
|
||||
\begin{tabular}{lc}
|
||||
\toprule
|
||||
\textbf{Benchmark} & \textbf{Score} \\
|
||||
\midrule
|
||||
VQA v2 & 88.5\% \\
|
||||
DesignBench & 82.4\% \\
|
||||
CLIP Score & 84.1\% \\
|
||||
FID Score & 73.5 \\
|
||||
\bottomrule
|
||||
\end{tabular}
|
||||
\caption{Visual Understanding Benchmarks}
|
||||
\end{table}
|
||||
|
||||
\subsection{Efficiency Metrics}
|
||||
|
||||
\begin{table}[H]
|
||||
\centering
|
||||
\begin{tabular}{ll}
|
||||
\toprule
|
||||
\textbf{Metric} & \textbf{Value} \\
|
||||
\midrule
|
||||
Inference Speed & 160 tokens/sec \\
|
||||
Memory Usage (INT4) & 4 GB \\
|
||||
Energy Efficiency & 93\% reduction \\
|
||||
Latency (First Token) & 50 ms \\
|
||||
\bottomrule
|
||||
\end{tabular}
|
||||
\caption{Efficiency Metrics}
|
||||
\end{table}
|
||||
|
||||
\section{Training Methodology}
|
||||
|
||||
\subsection{Dataset}
|
||||
The model was trained on a carefully curated dataset comprising:
|
||||
\begin{itemize}
|
||||
\item High-quality filtered web data (3TB)
|
||||
\item Domain-specific corpora for text-to-image generation
|
||||
\item Synthetic data generation for edge cases
|
||||
\item Human feedback through RLHF
|
||||
\end{itemize}
|
||||
|
||||
\subsection{Training Process}
|
||||
\begin{enumerate}
|
||||
\item \textbf{Pretraining}: 3 trillion tokens over 21 days on 16x A100
|
||||
\item \textbf{Supervised Fine-tuning}: Task-specific optimization
|
||||
\item \textbf{RLHF}: Alignment with human preferences
|
||||
\item \textbf{Constitutional AI}: Safety and helpfulness optimization
|
||||
\end{enumerate}
|
||||
|
||||
\section{Use Cases and Applications}
|
||||
|
||||
\subsection{Primary Applications}
|
||||
\item Creative content generation
|
||||
\item Marketing and advertising visuals
|
||||
\item Product design mockups
|
||||
\item Artistic style transfer
|
||||
\item Image restoration and enhancement
|
||||
|
||||
\subsection{Integration Examples}
|
||||
|
||||
\begin{lstlisting}[language=Python, caption=Basic Usage Example]
|
||||
from transformers import AutoModelForImageGeneration, AutoTokenizer
|
||||
|
||||
# Load model and tokenizer
|
||||
model = AutoModelForImageGeneration.from_pretrained("zenlm/zen-artist-8b")
|
||||
tokenizer = AutoTokenizer.from_pretrained("zenlm/zen-artist-8b")
|
||||
|
||||
# Generate response
|
||||
prompt = "A futuristic city at sunset"
|
||||
image = model.generate(prompt, num_inference_steps=50)
|
||||
image.save("generated_city.png")
|
||||
\end{lstlisting}
|
||||
|
||||
\section{Environmental Impact}
|
||||
|
||||
\subsection{Sustainability Metrics}
|
||||
\begin{itemize}
|
||||
\item \textbf{Carbon Footprint}: 0.09 kg CO₂e per million inferences
|
||||
\item \textbf{Energy Usage}: 2.0 kWh per day (1000 users)
|
||||
\item \textbf{Efficiency Gain}: 93\% reduction vs comparable models
|
||||
\end{itemize}
|
||||
|
||||
\subsection{Green AI Commitment}
|
||||
Zen AI models are designed with sustainability as a core principle, achieving industry-leading efficiency
|
||||
through architectural innovations and optimization techniques.
|
||||
|
||||
\section{Safety and Alignment}
|
||||
|
||||
\subsection{Safety Measures}
|
||||
\begin{itemize}
|
||||
\item Constitutional AI training for harmlessness
|
||||
\item Comprehensive red-teaming and adversarial testing
|
||||
\item Built-in safety filters and guardrails
|
||||
\item Regular safety audits and updates
|
||||
\end{itemize}
|
||||
|
||||
\subsection{Ethical Considerations}
|
||||
The model has been developed with careful attention to:
|
||||
\begin{itemize}
|
||||
\item Bias mitigation through diverse training data
|
||||
\item Transparency in capabilities and limitations
|
||||
\item Privacy-preserving deployment options
|
||||
\item Responsible AI principles alignment
|
||||
\end{itemize}
|
||||
|
||||
\section{Deployment Options}
|
||||
|
||||
\subsection{Available Formats}
|
||||
\begin{itemize}
|
||||
\item \textbf{SafeTensors}: Original precision weights
|
||||
\item \textbf{GGUF}: Quantized formats (Q4\_K\_M, Q5\_K\_M, Q8\_0)
|
||||
\item \textbf{MLX}: Apple Silicon optimization (4-bit, 8-bit)
|
||||
\item \textbf{ONNX}: Cross-platform deployment (coming soon)
|
||||
\end{itemize}
|
||||
|
||||
\subsection{Hardware Requirements}
|
||||
\begin{table}[H]
|
||||
\centering
|
||||
\begin{tabular}{lll}
|
||||
\toprule
|
||||
\textbf{Precision} & \textbf{Memory} & \textbf{Recommended Hardware} \\
|
||||
\midrule
|
||||
FP16 & 16 GB & RTX 3080 \\
|
||||
INT8 & 8 GB & RTX 3070 \\
|
||||
INT4 & 4 GB & M2 MacBook Air \\
|
||||
\bottomrule
|
||||
\end{tabular}
|
||||
\caption{Hardware Requirements by Precision}
|
||||
\end{table}
|
||||
|
||||
\section{Future Work}
|
||||
|
||||
\subsection{Planned Improvements}
|
||||
\begin{itemize}
|
||||
\item Extended context windows (up to 1M tokens)
|
||||
\item Enhanced multimodal capabilities
|
||||
\item Improved efficiency through further optimization
|
||||
\item Expanded language support
|
||||
\end{itemize}
|
||||
|
||||
\subsection{Research Directions}
|
||||
\begin{itemize}
|
||||
\item Advanced reasoning mechanisms
|
||||
\item Self-supervised learning improvements
|
||||
\item Zero-shot generalization enhancement
|
||||
\item Continual learning capabilities
|
||||
\end{itemize}
|
||||
|
||||
\section{Conclusion}
|
||||
|
||||
\textbf{Zen-Artist} represents a significant advancement in AI democratization,
|
||||
delivering exceptional performance for text-to-image generation while maintaining
|
||||
unprecedented efficiency. Through innovative architecture design and careful optimization,
|
||||
the model achieves a balance between capability and sustainability that sets a new standard
|
||||
for responsible AI development.
|
||||
|
||||
\section*{Acknowledgments}
|
||||
|
||||
We thank the open-source community, our research partners, and the teams at Hanzo AI and
|
||||
Zoo Labs Foundation for their contributions to this work.
|
||||
|
||||
\bibliographystyle{plain}
|
||||
\bibliography{references}
|
||||
|
||||
\appendix
|
||||
|
||||
\section{Model Card}
|
||||
|
||||
\begin{table}[H]
|
||||
\centering
|
||||
\begin{tabular}{ll}
|
||||
\toprule
|
||||
\textbf{Field} & \textbf{Value} \\
|
||||
\midrule
|
||||
Model Name & Zen-Artist \\
|
||||
Version & 1.0.0 \\
|
||||
Release Date & September 2025 \\
|
||||
License & Apache 2.0 \\
|
||||
Repository & \href{https://huggingface.co/zenlm/zen-artist-8b}{huggingface.co/zenlm/zen-artist-8b} \\
|
||||
Documentation & \href{https://github.com/zenlm/zen}{github.com/zenlm/zen} \\
|
||||
Contact & research@hanzo.ai \\
|
||||
\bottomrule
|
||||
\end{tabular}
|
||||
\caption{Model Card Information}
|
||||
\end{table}
|
||||
|
||||
\end{document}
|
||||
Binary file not shown.
@@ -2,6 +2,10 @@
|
||||
\usepackage[utf8]{inputenc}
|
||||
\usepackage[T1]{fontenc}
|
||||
\usepackage{amsmath,amsfonts,amssymb}
|
||||
\usepackage{amsthm}
|
||||
\newtheorem{theorem}{Theorem}
|
||||
\newtheorem{lemma}[theorem]{Lemma}
|
||||
\newtheorem{definition}[theorem]{Definition}
|
||||
\usepackage{graphicx}
|
||||
\usepackage{hyperref}
|
||||
\usepackage{listings}
|
||||
@@ -17,7 +21,7 @@
|
||||
|
||||
\title{\textbf{ASO: Active Semantic Optimization for Distributed AI}\\
|
||||
\large Technical Report v2025.05}
|
||||
\author{Zen LM Research Team\\
|
||||
\author{Antje Worring, Zach Kelling \\ Zen LM Research Team\\
|
||||
\texttt{research@zenlm.org}}
|
||||
\date{May 2025}
|
||||
|
||||
@@ -433,12 +437,12 @@ validation loss. ASO is available as an open protocol specification at
|
||||
codebase.
|
||||
|
||||
\section*{Acknowledgements}
|
||||
The Zen LM Research Team thanks the infrastructure team for cluster support and the
|
||||
The Antje Worring, Zach Kelling \\ Zen LM Research Team thanks the infrastructure team for cluster support and the
|
||||
evaluation team for benchmark maintenance.
|
||||
|
||||
\begin{thebibliography}{9}
|
||||
\bibitem{zenlm2025mode}
|
||||
Zen LM Research Team.
|
||||
Antje Worring, Zach Kelling \\ Zen LM Research Team.
|
||||
\textit{Zen MoDE: Mixture of Distilled Experts for Scalable Language Models}.
|
||||
Technical Report v2025.03, Zen LM, 2025.
|
||||
|
||||
|
||||
Binary file not shown.
@@ -15,7 +15,7 @@
|
||||
|
||||
\title{\textbf{Zen Audio Architecture: Universal Speech and Sound}\\
|
||||
\large Technical Report v2025.05}
|
||||
\author{Zen LM Research Team\\
|
||||
\author{Antje Worring, Zach Kelling \\ Zen LM Research Team\\
|
||||
\texttt{research@zenlm.org}}
|
||||
\date{May 2025}
|
||||
|
||||
|
||||
Binary file not shown.
@@ -16,7 +16,7 @@
|
||||
|
||||
\title{\textbf{Zen: A Foundation Language Model for Instruction Following}\\
|
||||
\large Technical Report v2025.01}
|
||||
\author{Zen LM Research Team\\
|
||||
\author{Antje Worring, Zach Kelling \\ Zen LM Research Team\\
|
||||
\texttt{research@zenlm.org}}
|
||||
\date{January 2025}
|
||||
|
||||
|
||||
Binary file not shown.
@@ -17,7 +17,7 @@
|
||||
\title{\textbf{ZenBench: A Comprehensive AI Evaluation Suite with\\
|
||||
Contamination Detection and Adaptive Calibration}\\
|
||||
\large Technical Report v2025.09}
|
||||
\author{Zen LM Research Team\\
|
||||
\author{Antje Worring, Zach Kelling \\ Zen LM Research Team\\
|
||||
\texttt{research@zenlm.org}}
|
||||
\date{September 2025}
|
||||
|
||||
|
||||
Binary file not shown.
@@ -16,7 +16,7 @@
|
||||
\title{\textbf{Chain-of-Thought Reasoning in Zen Models:\\
|
||||
Verified CoT Training with Step-Level Supervision}\\
|
||||
\large Technical Report v2025.05}
|
||||
\author{Zen LM Research Team\\
|
||||
\author{Antje Worring, Zach Kelling \\ Zen LM Research Team\\
|
||||
\texttt{research@zenlm.org}}
|
||||
\date{May 2025}
|
||||
|
||||
|
||||
Binary file not shown.
@@ -17,7 +17,7 @@
|
||||
|
||||
\title{\textbf{Comprehensive Code Intelligence Benchmarking for Zen}\\
|
||||
\large Technical Report v2025.09}
|
||||
\author{Zen LM Research Team\\
|
||||
\author{Antje Worring, Zach Kelling \\ Zen LM Research Team\\
|
||||
\texttt{research@zenlm.org}}
|
||||
\date{September 2025}
|
||||
|
||||
|
||||
@@ -1,476 +0,0 @@
|
||||
\documentclass[11pt,a4paper]{article}
|
||||
\usepackage[utf8]{inputenc}
|
||||
\usepackage[T1]{fontenc}
|
||||
\usepackage{amsmath,amsfonts,amssymb}
|
||||
\usepackage{graphicx}
|
||||
\usepackage{hyperref}
|
||||
\usepackage{listings}
|
||||
\usepackage{color}
|
||||
\usepackage{booktabs}
|
||||
\usepackage{float}
|
||||
\usepackage{geometry}
|
||||
\geometry{margin=1in}
|
||||
\definecolor{zenblue}{RGB}{41,121,255}
|
||||
\definecolor{zengreen}{RGB}{52,199,89}
|
||||
\definecolor{codegray}{RGB}{240,240,240}
|
||||
\hypersetup{colorlinks=true,linkcolor=zenblue,urlcolor=zenblue,citecolor=zenblue}
|
||||
|
||||
\lstset{
|
||||
backgroundcolor=\color{codegray},
|
||||
basicstyle=\ttfamily\small,
|
||||
breaklines=true,
|
||||
frame=single,
|
||||
language=Python
|
||||
}
|
||||
|
||||
\title{\textbf{Zen-Code: Code Completion and Intelligence at 14B Scale}\\
|
||||
\large Technical Report v2024.12}
|
||||
\author{Zen LM Research Team\\
|
||||
\texttt{research@zenlm.org}}
|
||||
\date{December 2024}
|
||||
|
||||
\begin{document}
|
||||
\maketitle
|
||||
|
||||
\begin{abstract}
|
||||
We present \textbf{Zen-Code}, a 14 billion parameter language model optimized for code
|
||||
generation, completion, and understanding across 80$+$ programming languages. Built on the
|
||||
Zen MoDE (Mixture of Distilled Experts) architecture, Zen-Code is pretrained on 2T tokens of
|
||||
high-quality code with fill-in-the-middle (FIM) objectives and fine-tuned on repository-scale
|
||||
context tasks. Zen-Code achieves HumanEval 87.2\%, MBPP 82.3\%, SWE-bench 28.4\%, and
|
||||
RepoBench 0.834, establishing it as the capable code-specialist in the Zen family. The model
|
||||
supports long repository-level context up to 64K tokens, multi-file awareness, and standard
|
||||
IDE integration formats including LSP-compatible completions.
|
||||
\end{abstract}
|
||||
|
||||
\tableofcontents
|
||||
\newpage
|
||||
|
||||
%% ─────────────────────────────────────────────────────────────────────────────
|
||||
\section{Introduction}
|
||||
|
||||
Code generation has emerged as one of the most commercially significant applications of large
|
||||
language models. From IDE autocomplete to automated pull request generation, models that reliably
|
||||
understand and generate code at the repository scale provide substantial productivity gains
|
||||
\cite{chen2021codex, nijkamp2023codegen2}. General-purpose language models, while capable of
|
||||
code generation, are not optimized for the specific properties of software: multi-file context,
|
||||
precise syntax requirements, test-driven correctness, and sub-token completion granularity.
|
||||
|
||||
Zen-Code addresses this gap with a model purpose-built for code. Key contributions:
|
||||
|
||||
\begin{itemize}
|
||||
\item A 14B parameter model pretrained on 2T code tokens from 80$+$ languages, with an
|
||||
additional 500B tokens of technical documentation and Stack Exchange content.
|
||||
\item Fill-in-the-middle (FIM) training enabling prefix-suffix-middle completion, the
|
||||
format required by production IDE integrations.
|
||||
\item Repository-level context training on 64K token windows with structured file-path
|
||||
and symbol-table prefixes.
|
||||
\item Post-training on instruction-following code tasks including debugging, refactoring,
|
||||
test generation, and code review.
|
||||
\item Benchmark results: HumanEval 87.2\%, MBPP 82.3\%, SWE-bench 28.4\%, RepoBench 0.834.
|
||||
\end{itemize}
|
||||
|
||||
%% ─────────────────────────────────────────────────────────────────────────────
|
||||
\section{Architecture}
|
||||
|
||||
\subsection{Model Configuration}
|
||||
|
||||
Zen-Code uses the Zen MoDE architecture scaled to 14B parameters, a tier chosen to balance
|
||||
code reasoning capability with inference cost on developer hardware (single A100 or 2$\times$
|
||||
RTX 4090 in BF16).
|
||||
|
||||
\begin{table}[H]
|
||||
\centering
|
||||
\caption{Zen-Code architecture hyperparameters.}
|
||||
\label{tab:arch}
|
||||
\begin{tabular}{lc}
|
||||
\toprule
|
||||
\textbf{Hyperparameter} & \textbf{Value} \\
|
||||
\midrule
|
||||
Parameters (total) & 14.4B \\
|
||||
Layers & 40 \\
|
||||
Attention heads & 40 \\
|
||||
KV heads (GQA) & 8 \\
|
||||
Hidden dimension & 5120 \\
|
||||
FFN intermediate dimension & 13{,}696 \\
|
||||
Vocabulary size & 151{,}936 \\
|
||||
Context length (training) & 65{,}536 \\
|
||||
Position encoding & RoPE ($\theta = 2{,}000{,}000$) \\
|
||||
Activation & SiLU \\
|
||||
Normalization & RMSNorm \\
|
||||
FIM tokens & \texttt{<|fim\_prefix|>}, \texttt{<|fim\_middle|>}, \texttt{<|fim\_suffix|>} \\
|
||||
\bottomrule
|
||||
\end{tabular}
|
||||
\end{table}
|
||||
|
||||
\subsection{Fill-in-the-Middle Token Scheme}
|
||||
|
||||
Zen-Code supports fill-in-the-middle (FIM) completion \cite{bavarian2022fim}, the standard
|
||||
format for IDE inline completion. The three FIM sentinel tokens partition the input:
|
||||
|
||||
\begin{lstlisting}[language={},caption={FIM token format.}]
|
||||
<|fim_prefix|>[prefix code]<|fim_suffix|>[suffix code]<|fim_middle|>
|
||||
\end{lstlisting}
|
||||
|
||||
During training, 50\% of code samples are converted to FIM format using the PSM
|
||||
(prefix-suffix-middle) or SPM (suffix-prefix-middle) transformation drawn with equal
|
||||
probability, following the recommendation of \cite{bavarian2022fim}.
|
||||
|
||||
For the PSM transformation of a document $d$ split at position $s$:
|
||||
|
||||
\begin{equation}
|
||||
d_{\text{FIM}} = [\texttt{FP}] \cdot d_{[:s]} \cdot [\texttt{FS}] \cdot d_{[s:]} \cdot [\texttt{FM}] \cdot d_{\text{middle}}
|
||||
\end{equation}
|
||||
|
||||
where $\texttt{FP}, \texttt{FS}, \texttt{FM}$ are the prefix, suffix, and middle sentinels
|
||||
respectively, and $d_{\text{middle}}$ is the held-out span the model must predict.
|
||||
|
||||
\subsection{Repository-Level Context Structure}
|
||||
|
||||
For long-context code tasks, Zen-Code accepts structured repository context using a
|
||||
lightweight file-header format:
|
||||
|
||||
\begin{lstlisting}[language={},caption={Repository context format.}]
|
||||
# File: src/utils/parser.py
|
||||
[file content]
|
||||
|
||||
# File: src/models/base.py
|
||||
[file content]
|
||||
|
||||
# File: [target file, completion requested here]
|
||||
[prefix]<|fim_suffix|>[suffix]<|fim_middle|>
|
||||
\end{lstlisting}
|
||||
|
||||
This format is injected by IDE extensions and agents to provide cross-file context.
|
||||
The 64K context window accommodates typical repository relevant-file sets (10--50 files
|
||||
of 500--1000 lines each).
|
||||
|
||||
%% ─────────────────────────────────────────────────────────────────────────────
|
||||
\section{Training Methodology}
|
||||
|
||||
\subsection{Pretraining Data}
|
||||
|
||||
Zen-Code's 2.5T-token pretraining corpus is heavily code-dominated.
|
||||
|
||||
\begin{table}[H]
|
||||
\centering
|
||||
\caption{Zen-Code pretraining data composition.}
|
||||
\label{tab:data}
|
||||
\begin{tabular}{lcc}
|
||||
\toprule
|
||||
\textbf{Domain} & \textbf{Tokens (B)} & \textbf{Fraction} \\
|
||||
\midrule
|
||||
Source code (80$+$ languages) & 1{,}500 & 60.0\% \\
|
||||
Technical documentation & 375 & 15.0\% \\
|
||||
Stack Overflow / Q\&A & 250 & 10.0\% \\
|
||||
Research papers (CS/ML) & 125 & 5.0\% \\
|
||||
General web (filtered) & 125 & 5.0\% \\
|
||||
Formal specs (Lean, Coq, TLA$+$) & 75 & 3.0\% \\
|
||||
Synthetic code (unit tests) & 50 & 2.0\% \\
|
||||
\midrule
|
||||
Total & 2{,}500 & 100.0\% \\
|
||||
\bottomrule
|
||||
\end{tabular}
|
||||
\end{table}
|
||||
|
||||
\subsection{Per-Language Distribution}
|
||||
|
||||
\begin{table}[H]
|
||||
\centering
|
||||
\caption{Top programming languages by token count in pretraining corpus.}
|
||||
\label{tab:langs}
|
||||
\begin{tabular}{lcc}
|
||||
\toprule
|
||||
\textbf{Language} & \textbf{Tokens (B)} & \textbf{Share of code data} \\
|
||||
\midrule
|
||||
Python & 360 & 24.0\% \\
|
||||
JavaScript & 240 & 16.0\% \\
|
||||
TypeScript & 150 & 10.0\% \\
|
||||
Java & 135 & 9.0\% \\
|
||||
C/C$+$$+$ & 120 & 8.0\% \\
|
||||
Rust & 90 & 6.0\% \\
|
||||
Go & 75 & 5.0\% \\
|
||||
Ruby & 45 & 3.0\% \\
|
||||
PHP & 45 & 3.0\% \\
|
||||
Shell/Bash & 45 & 3.0\% \\
|
||||
Other (70$+$) & 195 & 13.0\% \\
|
||||
\midrule
|
||||
Total & 1{,}500 & 100.0\% \\
|
||||
\bottomrule
|
||||
\end{tabular}
|
||||
\end{table}
|
||||
|
||||
Data deduplication uses a two-stage process: exact SHA-256 deduplication at the file level,
|
||||
followed by MinHash near-deduplication with 10-shingling at LSH threshold 0.8 to remove
|
||||
forked repositories and boilerplate.
|
||||
|
||||
\subsection{Pretraining Procedure}
|
||||
|
||||
Zen-Code is trained from scratch (not initialized from Zen base) to allow the training
|
||||
distribution to be fully code-optimized without general text forgetting constraints.
|
||||
|
||||
\begin{itemize}
|
||||
\item Optimizer: AdamW, LR $3 \times 10^{-4}$ $\to$ $3 \times 10^{-5}$ (cosine)
|
||||
\item Warm-up: 1000 steps
|
||||
\item Weight decay: 0.1
|
||||
\item Batch size: 4M tokens
|
||||
\item FIM rate: 50\% of code samples
|
||||
\item Precision: BF16
|
||||
\item Training duration: $\approx$625K steps, $\approx$512 H100 GPUs
|
||||
\end{itemize}
|
||||
|
||||
\subsection{Repository-Level Fine-Tuning}
|
||||
|
||||
After base pretraining, Zen-Code undergoes a 50B-token repository-level fine-tuning phase
|
||||
using full repository snapshots. Repositories are sampled from GitHub with permissive licenses,
|
||||
filtered for non-trivial size (>10 files, >1K lines). Files within each repository are packed
|
||||
into 64K-token windows following a topological ordering based on import dependencies.
|
||||
|
||||
\subsection{Instruction Fine-Tuning}
|
||||
|
||||
A final SFT phase on 800K instruction-code pairs covering:
|
||||
|
||||
\begin{itemize}
|
||||
\item Code generation from natural language descriptions
|
||||
\item Bug finding and fixing (with error message context)
|
||||
\item Code explanation and documentation generation
|
||||
\item Unit test generation from function signatures
|
||||
\item Code refactoring and style improvement
|
||||
\item API usage from documentation
|
||||
\end{itemize}
|
||||
|
||||
SFT learning rate: $1 \times 10^{-5}$, 2 epochs.
|
||||
|
||||
%% ─────────────────────────────────────────────────────────────────────────────
|
||||
\section{Evaluation}
|
||||
|
||||
\subsection{Code Generation Benchmarks}
|
||||
|
||||
\begin{table}[H]
|
||||
\centering
|
||||
\caption{Code generation benchmark results.}
|
||||
\label{tab:benchmarks}
|
||||
\begin{tabular}{lcccc}
|
||||
\toprule
|
||||
\textbf{Benchmark} & \textbf{Zen-Code (14B)} & \textbf{Comp.\ A (15B)} & \textbf{Comp.\ B (13B)} & \textbf{Comp.\ C (7B)} \\
|
||||
\midrule
|
||||
HumanEval (pass@1) & \textbf{87.2} & 85.1 & 83.6 & 76.3 \\
|
||||
HumanEval$+$ (pass@1) & \textbf{82.4} & 80.3 & 78.7 & 71.2 \\
|
||||
MBPP (pass@1) & \textbf{82.3} & 80.1 & 78.4 & 70.8 \\
|
||||
MBPP$+$ (pass@1) & \textbf{76.8} & 74.3 & 72.1 & 64.3 \\
|
||||
LiveCodeBench (3-month) & \textbf{54.2} & 51.7 & 49.3 & 41.8 \\
|
||||
\midrule
|
||||
SWE-bench Verified & \textbf{28.4} & 25.3 & 22.8 & 14.6 \\
|
||||
RepoBench (retrieval) & \textbf{0.834} & 0.812 & 0.798 & 0.741 \\
|
||||
\midrule
|
||||
MultiPL-E (avg 12 langs)& \textbf{72.4} & 70.8 & 68.3 & 61.2 \\
|
||||
\bottomrule
|
||||
\end{tabular}
|
||||
\end{table}
|
||||
|
||||
\subsection{Multilingual Code Evaluation}
|
||||
|
||||
\begin{table}[H]
|
||||
\centering
|
||||
\caption{MultiPL-E pass@1 by programming language.}
|
||||
\label{tab:multilingual_code}
|
||||
\begin{tabular}{lcc}
|
||||
\toprule
|
||||
\textbf{Language} & \textbf{Zen-Code (14B)} & \textbf{Competitor (15B)} \\
|
||||
\midrule
|
||||
Python & 87.2 & 85.1 \\
|
||||
JavaScript & 78.6 & 76.4 \\
|
||||
TypeScript & 76.4 & 74.1 \\
|
||||
Java & 74.3 & 72.8 \\
|
||||
C$+$$+$ & 70.8 & 68.3 \\
|
||||
Rust & 68.4 & 65.7 \\
|
||||
Go & 73.1 & 71.4 \\
|
||||
Ruby & 65.3 & 62.8 \\
|
||||
PHP & 63.7 & 61.4 \\
|
||||
Shell & 58.4 & 55.9 \\
|
||||
\bottomrule
|
||||
\end{tabular}
|
||||
\end{table}
|
||||
|
||||
\subsection{Fill-in-the-Middle Evaluation}
|
||||
|
||||
\begin{table}[H]
|
||||
\centering
|
||||
\caption{FIM completion accuracy on HumanEval-Infilling.}
|
||||
\label{tab:fim}
|
||||
\begin{tabular}{lcc}
|
||||
\toprule
|
||||
\textbf{Metric} & \textbf{Zen-Code (14B)} & \textbf{Competitor (15B)} \\
|
||||
\midrule
|
||||
Single-line (exact match) & 82.3 & 79.1 \\
|
||||
Multi-line (exact match) & 61.4 & 57.8 \\
|
||||
Single-line (functional) & 91.7 & 89.4 \\
|
||||
Multi-line (functional) & 74.6 & 71.3 \\
|
||||
\bottomrule
|
||||
\end{tabular}
|
||||
\end{table}
|
||||
|
||||
\subsection{Repository-Level Evaluation}
|
||||
|
||||
SWE-bench Verified requires understanding a full repository, identifying which files to
|
||||
modify, and generating a correct patch that passes the repository's test suite.
|
||||
|
||||
\begin{table}[H]
|
||||
\centering
|
||||
\caption{SWE-bench Verified resolution rates by issue category.}
|
||||
\label{tab:swebench}
|
||||
\begin{tabular}{lc}
|
||||
\toprule
|
||||
\textbf{Category} & \textbf{Zen-Code (14B)} \\
|
||||
\midrule
|
||||
Logic errors & 34.2\% \\
|
||||
Type errors & 41.8\% \\
|
||||
Off-by-one & 38.7\% \\
|
||||
Missing handling & 31.4\% \\
|
||||
API usage errors & 29.6\% \\
|
||||
Regression fixes & 18.3\% \\
|
||||
\midrule
|
||||
Overall (verified) & 28.4\% \\
|
||||
\bottomrule
|
||||
\end{tabular}
|
||||
\end{table}
|
||||
|
||||
%% ─────────────────────────────────────────────────────────────────────────────
|
||||
\section{IDE Integration}
|
||||
|
||||
Zen-Code is designed for IDE integration via a thin inference server exposing an
|
||||
OpenAI-compatible completions API with FIM support:
|
||||
|
||||
\begin{lstlisting}[language=Python, caption={Example FIM completion request.}]
|
||||
import openai
|
||||
|
||||
client = openai.OpenAI(
|
||||
base_url="https://api.zenlm.org/v1",
|
||||
api_key="<zen-api-key>"
|
||||
)
|
||||
|
||||
response = client.completions.create(
|
||||
model="zen-code-14b",
|
||||
prompt="<|fim_prefix|>def merge_sorted_arrays(a, b):\n "
|
||||
"<|fim_suffix|>\n return result\n<|fim_middle|>",
|
||||
max_tokens=128,
|
||||
temperature=0.1,
|
||||
stop=["<|endoftext|>"]
|
||||
)
|
||||
|
||||
print(response.choices[0].text)
|
||||
\end{lstlisting}
|
||||
|
||||
Available IDE integrations include VS Code (via Continue and Tabby plugins), JetBrains
|
||||
IDEs, Neovim (via coc.nvim), and Emacs (via copilot.el-compatible backends).
|
||||
|
||||
%% ─────────────────────────────────────────────────────────────────────────────
|
||||
\section{Ablation Studies}
|
||||
|
||||
\subsection{Effect of FIM Training Rate}
|
||||
|
||||
\begin{table}[H]
|
||||
\centering
|
||||
\caption{HumanEval-Infilling (single-line) vs.\ FIM training fraction.}
|
||||
\label{tab:fim_ablation}
|
||||
\begin{tabular}{lcc}
|
||||
\toprule
|
||||
\textbf{FIM rate} & \textbf{HumanEval pass@1} & \textbf{FIM exact match (single)} \\
|
||||
\midrule
|
||||
0\% (no FIM) & 86.8 & 41.2 \\
|
||||
25\% & 85.9 & 74.3 \\
|
||||
50\% & 87.2 & 82.3 \\
|
||||
75\% & 84.3 & 83.1 \\
|
||||
\bottomrule
|
||||
\end{tabular}
|
||||
\end{table}
|
||||
|
||||
50\% FIM rate provides the optimal balance: full left-to-right completion performance is
|
||||
maintained (87.2\% vs.\ 86.8\% baseline) while infilling performance improves dramatically.
|
||||
At 75\% FIM the model under-trains on left-to-right generation, reducing HumanEval by 2.9pp.
|
||||
|
||||
\subsection{Effect of Repository Fine-Tuning}
|
||||
|
||||
\begin{table}[H]
|
||||
\centering
|
||||
\caption{Impact of repository-level fine-tuning on long-context tasks.}
|
||||
\label{tab:repo_ablation}
|
||||
\begin{tabular}{lcc}
|
||||
\toprule
|
||||
\textbf{Model} & \textbf{RepoBench} & \textbf{SWE-bench Verified} \\
|
||||
\midrule
|
||||
Base (no repo FT) & 0.761 & 19.3 \\
|
||||
$+$ Repo FT (8K context) & 0.803 & 23.7 \\
|
||||
$+$ Repo FT (64K context) & 0.834 & 28.4 \\
|
||||
\bottomrule
|
||||
\end{tabular}
|
||||
\end{table}
|
||||
|
||||
%% ─────────────────────────────────────────────────────────────────────────────
|
||||
\section{Related Work}
|
||||
|
||||
Code-specialized language models trace to Codex \cite{chen2021codex}, which demonstrated
|
||||
that training on large code corpora enables competitive function-level generation. CodeGen
|
||||
\cite{nijkamp2023codegen2} and its successors explored scaling laws for code. StarCoder
|
||||
\cite{li2023starcoder} trained on The Stack, a permissively licensed code corpus with strong
|
||||
deduplication. Fill-in-the-middle training was introduced in \cite{bavarian2022fim} and
|
||||
adopted by most subsequent code models. SWE-bench \cite{jimenez2023swe} introduced
|
||||
repository-level task solving as a benchmark; the verified subset reduces noise in evaluation.
|
||||
RepoBench \cite{liu2023repobench} tests long-context retrieval within repository structure.
|
||||
|
||||
%% ─────────────────────────────────────────────────────────────────────────────
|
||||
\section{Limitations}
|
||||
|
||||
Zen-Code's SWE-bench performance of 28.4\% indicates that most real-world repository-level
|
||||
issues remain unsolved in single-pass generation without agentic scaffolding. The model does
|
||||
not have access to runtime execution during generation; code correctness is assessed only at
|
||||
test time. Languages with fewer than 10B training tokens show substantially weaker performance
|
||||
than top-tier languages. Security-sensitive code generation (cryptography, access control)
|
||||
should be reviewed by experts regardless of benchmark scores.
|
||||
|
||||
%% ─────────────────────────────────────────────────────────────────────────────
|
||||
\section{Conclusion}
|
||||
|
||||
Zen-Code provides a purpose-built code intelligence model at 14B parameters, delivering
|
||||
HumanEval 87.2\%, MBPP 82.3\%, SWE-bench 28.4\%, and RepoBench 0.834 through a combination
|
||||
of code-dominated pretraining, fill-in-the-middle objectives, and repository-scale context
|
||||
fine-tuning. The model's 64K context window, FIM support, and IDE-compatible API make it
|
||||
a practical choice for developer tooling deployment. Zen-Code is the foundational code model
|
||||
in the Zen family, with Zen-Coder-Flash providing a distilled fast-inference variant for
|
||||
latency-critical IDE autocomplete applications.
|
||||
|
||||
%% ─────────────────────────────────────────────────────────────────────────────
|
||||
\begin{thebibliography}{99}
|
||||
|
||||
\bibitem{chen2021codex}
|
||||
M.~Chen et al., ``Evaluating Large Language Models Trained on Code,''
|
||||
\textit{arXiv:2107.03374}, 2021.
|
||||
|
||||
\bibitem{nijkamp2023codegen2}
|
||||
E.~Nijkamp et al., ``CodeGen2: Lessons for Training LLMs on Programming and Natural Languages,''
|
||||
\textit{ICLR}, 2023.
|
||||
|
||||
\bibitem{bavarian2022fim}
|
||||
M.~Bavarian et al., ``Efficient Training of Language Models to Fill in the Middle,''
|
||||
\textit{arXiv:2207.14255}, 2022.
|
||||
|
||||
\bibitem{li2023starcoder}
|
||||
R.~Li et al., ``StarCoder: May the Source Be with You!'' \textit{arXiv:2305.06161}, 2023.
|
||||
|
||||
\bibitem{jimenez2023swe}
|
||||
C.~Jimenez et al., ``SWE-bench: Can Language Models Resolve Real-World GitHub Issues?''
|
||||
\textit{arXiv:2310.06770}, 2023.
|
||||
|
||||
\bibitem{liu2023repobench}
|
||||
T.~Liu et al., ``RepoBench: Benchmarking Repository-Level Code Auto-Completion Systems,''
|
||||
\textit{arXiv:2306.03091}, 2023.
|
||||
|
||||
\bibitem{loshchilov2019decoupled}
|
||||
I.~Loshchilov and F.~Hutter, ``Decoupled Weight Decay Regularization,'' \textit{ICLR}, 2019.
|
||||
|
||||
\bibitem{broder1997minwise}
|
||||
A.~Broder, ``On the resemblance and containment of documents,'' \textit{Sequences}, 1997.
|
||||
|
||||
\end{thebibliography}
|
||||
|
||||
\end{document}
|
||||
Binary file not shown.
@@ -24,7 +24,7 @@
|
||||
|
||||
\title{\textbf{Zen-Coder-Flash: Ultra-Low-Latency Code Completion via Knowledge Distillation}\\
|
||||
\large Technical Report v2025.02}
|
||||
\author{Zen LM Research Team\\
|
||||
\author{Antje Worring, Zach Kelling \\ Zen LM Research Team\\
|
||||
\texttt{research@zenlm.org}}
|
||||
\date{February 2025}
|
||||
|
||||
@@ -211,7 +211,7 @@ Feature coefficient $\gamma$ & 0.2 \\
|
||||
Duration & 500B tokens (250K steps) \\
|
||||
\bottomrule
|
||||
\end{tabular}
|
||||
\end{table>}
|
||||
\end{table}
|
||||
|
||||
\subsection{Speculative Decoding Setup}
|
||||
|
||||
|
||||
Binary file not shown.
@@ -5,7 +5,7 @@
|
||||
\usepackage{graphicx}
|
||||
\usepackage{hyperref}
|
||||
\usepackage{listings}
|
||||
\usepackage{color}
|
||||
\usepackage[dvipsnames]{xcolor}
|
||||
\usepackage{booktabs}
|
||||
\usepackage{float}
|
||||
\usepackage{geometry}
|
||||
@@ -48,7 +48,7 @@
|
||||
}
|
||||
|
||||
\author{
|
||||
Zach Kelling\thanks{zach@lux.network} \\
|
||||
Antje Worring, Zach Kelling\thanks{zach@lux.network} \\
|
||||
\texttt{research@hanzo.ai} \\
|
||||
\\
|
||||
Zoo Labs Foundation \\
|
||||
|
||||
Binary file not shown.
@@ -17,7 +17,7 @@
|
||||
|
||||
\title{\textbf{Long Context Scaling in Zen Models: 1M Token Extension}\\
|
||||
\large Technical Report v2025.07}
|
||||
\author{Zen LM Research Team\\
|
||||
\author{Antje Worring, Zach Kelling \\ Zen LM Research Team\\
|
||||
\texttt{research@zenlm.org}}
|
||||
\date{July 2025}
|
||||
|
||||
|
||||
Binary file not shown.
@@ -5,7 +5,7 @@
|
||||
\usepackage{graphicx}
|
||||
\usepackage{hyperref}
|
||||
\usepackage{listings}
|
||||
\usepackage{color}
|
||||
\usepackage{xcolor}
|
||||
\usepackage{booktabs}
|
||||
\usepackage{float}
|
||||
\usepackage{geometry}
|
||||
@@ -48,7 +48,7 @@
|
||||
}
|
||||
|
||||
\author{
|
||||
Zach Kelling\thanks{zach@lux.network} \\
|
||||
Antje Worring, Zach Kelling\thanks{zach@lux.network} \\
|
||||
\texttt{research@hanzo.ai} \\
|
||||
\\
|
||||
Zoo Labs Foundation \\
|
||||
@@ -77,13 +77,14 @@ The rapid advancement of artificial intelligence has created an unprecedented de
|
||||
|
||||
\subsection{Key Innovations}
|
||||
\begin{itemize}
|
||||
egin{itemize}
|
||||
\item \textbf{Efficient Architecture}: 22B active parameters from 235B total
|
||||
\item \textbf{Specialized Training}: Optimized for design generation
|
||||
\item \textbf{Extended Context}: 131K context window
|
||||
\item \textbf{Thinking Mode}: 512K thinking tokens
|
||||
|
||||
|
||||
\end{itemize}
|
||||
|
||||
|
||||
|
||||
\section{Architecture}
|
||||
|
||||
@@ -121,6 +122,8 @@ Specialized attention mechanisms optimized for design generation.
|
||||
|
||||
\subsubsection{Thinking Mode}
|
||||
Advanced reasoning through extended thinking tokens (up to 512K), enabling:
|
||||
\begin{itemize}
|
||||
egin{itemize}
|
||||
\begin{itemize}
|
||||
\item Step-by-step problem decomposition
|
||||
\item Self-correction and verification
|
||||
@@ -170,23 +173,25 @@ Latency (First Token) & 180 ms \\
|
||||
\subsection{Dataset}
|
||||
The model was trained on a carefully curated dataset comprising:
|
||||
\begin{itemize}
|
||||
egin{itemize}
|
||||
\item High-quality filtered web data (50TB)
|
||||
\item Domain-specific corpora for design generation
|
||||
\item Synthetic data generation for edge cases
|
||||
\item Human feedback through RLHF
|
||||
\end{itemize}
|
||||
|
||||
\subsection{Training Process}
|
||||
\begin{enumerate}
|
||||
egin{itemize}
|
||||
\item \textbf{Pretraining}: 7 trillion tokens over 60 days on 128x A100
|
||||
\item \textbf{Supervised Fine-tuning}: Task-specific optimization
|
||||
\item \textbf{RLHF}: Alignment with human preferences
|
||||
\item \textbf{Constitutional AI}: Safety and helpfulness optimization
|
||||
\end{enumerate}
|
||||
|
||||
\section{Use Cases and Applications}
|
||||
|
||||
\subsection{Primary Applications}
|
||||
\begin{itemize}
|
||||
egin{itemize}
|
||||
\item UI/UX design analysis
|
||||
\item Architecture and layout planning
|
||||
\item Visual question answering
|
||||
@@ -212,10 +217,10 @@ analysis = processor.decode(outputs[0])
|
||||
|
||||
\subsection{Sustainability Metrics}
|
||||
\begin{itemize}
|
||||
\item \textbf{Carbon Footprint}: 0.35 kg CO₂e per million inferences
|
||||
egin{itemize}
|
||||
\item \textbf{Carbon Footprint}: 0.35 kg CO\textsubscript{2}e per million inferences
|
||||
\item \textbf{Energy Usage}: 8.0 kWh per day (1000 users)
|
||||
\item \textbf{Efficiency Gain}: 90\% reduction vs comparable models
|
||||
\end{itemize}
|
||||
|
||||
\subsection{Green AI Commitment}
|
||||
Zen AI models are designed with sustainability as a core principle, achieving industry-leading efficiency
|
||||
@@ -225,30 +230,30 @@ through architectural innovations and optimization techniques.
|
||||
|
||||
\subsection{Safety Measures}
|
||||
\begin{itemize}
|
||||
egin{itemize}
|
||||
\item Constitutional AI training for harmlessness
|
||||
\item Comprehensive red-teaming and adversarial testing
|
||||
\item Built-in safety filters and guardrails
|
||||
\item Regular safety audits and updates
|
||||
\end{itemize}
|
||||
|
||||
\subsection{Ethical Considerations}
|
||||
The model has been developed with careful attention to:
|
||||
\begin{itemize}
|
||||
egin{itemize}
|
||||
\item Bias mitigation through diverse training data
|
||||
\item Transparency in capabilities and limitations
|
||||
\item Privacy-preserving deployment options
|
||||
\item Responsible AI principles alignment
|
||||
\end{itemize}
|
||||
|
||||
\section{Deployment Options}
|
||||
|
||||
\subsection{Available Formats}
|
||||
\begin{itemize}
|
||||
egin{itemize}
|
||||
\item \textbf{SafeTensors}: Original precision weights
|
||||
\item \textbf{GGUF}: Quantized formats (Q4\_K\_M, Q5\_K\_M, Q8\_0)
|
||||
\item \textbf{MLX}: Apple Silicon optimization (4-bit, 8-bit)
|
||||
\item \textbf{ONNX}: Cross-platform deployment (coming soon)
|
||||
\end{itemize}
|
||||
|
||||
\subsection{Hardware Requirements}
|
||||
\begin{table}[H]
|
||||
@@ -269,19 +274,19 @@ INT4 & 55 GB & A100 80GB \\
|
||||
|
||||
\subsection{Planned Improvements}
|
||||
\begin{itemize}
|
||||
egin{itemize}
|
||||
\item Extended context windows (up to 1M tokens)
|
||||
\item Enhanced multimodal capabilities
|
||||
\item Improved efficiency through further optimization
|
||||
\item Expanded language support
|
||||
\end{itemize}
|
||||
|
||||
\subsection{Research Directions}
|
||||
\begin{itemize}
|
||||
egin{itemize}
|
||||
\item Advanced reasoning mechanisms
|
||||
\item Self-supervised learning improvements
|
||||
\item Zero-shot generalization enhancement
|
||||
\item Continual learning capabilities
|
||||
\end{itemize}
|
||||
|
||||
\section{Conclusion}
|
||||
|
||||
@@ -296,8 +301,19 @@ for responsible AI development.
|
||||
We thank the open-source community, our research partners, and the teams at Hanzo AI and
|
||||
Zoo Labs Foundation for their contributions to this work.
|
||||
|
||||
\bibliographystyle{plain}
|
||||
\bibliography{references}
|
||||
\begin{thebibliography}{99}
|
||||
\bibitem{vaswani2017attention} Vaswani, A. et al. (2017). Attention Is All You Need. NeurIPS 2017.
|
||||
\bibitem{brown2020language} Brown, T. et al. (2020). Language Models are Few-Shot Learners. NeurIPS 2020.
|
||||
\bibitem{ouyang2022training} Ouyang, L. et al. (2022). Training Language Models to Follow Instructions with Human Feedback. NeurIPS 2022.
|
||||
\bibitem{ho2020denoising} Ho, J., Jain, A. and Abbeel, P. (2020). Denoising Diffusion Probabilistic Models. NeurIPS 2020.
|
||||
\bibitem{rombach2022high} Rombach, R. et al. (2022). High-Resolution Image Synthesis with Latent Diffusion Models. CVPR 2022.
|
||||
\bibitem{radford2021learning} Radford, A. et al. (2021). Learning Transferable Visual Models From Natural Language Supervision. ICML 2021.
|
||||
\bibitem{shazeer2017outrageously} Shazeer, N. et al. (2017). Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer. ICLR 2017.
|
||||
\bibitem{fedus2022switch} Fedus, W. et al. (2022). Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity. JMLR 2022.
|
||||
\bibitem{rafailov2023direct} Rafailov, R. et al. (2023). Direct Preference Optimization: Your Language Model is Secretly a Reward Model. NeurIPS 2023.
|
||||
\bibitem{schulman2017proximal} Schulman, J. et al. (2017). Proximal Policy Optimization Algorithms. arXiv:1707.06347.
|
||||
\bibitem{touvron2023llama} Touvron, H. et al. (2023). LLaMA: Open and Efficient Foundation Language Models. arXiv:2302.13971.
|
||||
\end{thebibliography}
|
||||
|
||||
\appendix
|
||||
|
||||
|
||||
Binary file not shown.
@@ -5,7 +5,7 @@
|
||||
\usepackage{graphicx}
|
||||
\usepackage{hyperref}
|
||||
\usepackage{listings}
|
||||
\usepackage{color}
|
||||
\usepackage{xcolor}
|
||||
\usepackage{booktabs}
|
||||
\usepackage{float}
|
||||
\usepackage{geometry}
|
||||
@@ -42,13 +42,13 @@
|
||||
\vspace{0.5cm}
|
||||
\Huge \textbf{Zen-Designer-Thinking} \\
|
||||
\vspace{0.3cm}
|
||||
\large Visual Reasoning & Analysis \\
|
||||
\large Visual Reasoning \& Analysis \\
|
||||
\vspace{0.5cm}
|
||||
\normalsize Technical Whitepaper v1.0
|
||||
}
|
||||
|
||||
\author{
|
||||
Zach Kelling\thanks{zach@lux.network} \\
|
||||
Antje Worring, Zach Kelling\thanks{zach@lux.network} \\
|
||||
\texttt{research@hanzo.ai} \\
|
||||
\\
|
||||
Zoo Labs Foundation \\
|
||||
@@ -62,7 +62,7 @@
|
||||
\maketitle
|
||||
|
||||
\begin{abstract}
|
||||
We present \textbf{Zen-Designer-Thinking}, a 235B parameter model optimized for visual reasoning & analysis.
|
||||
We present \textbf{Zen-Designer-Thinking}, a 235B parameter model optimized for visual reasoning \& analysis.
|
||||
Built upon a frontier vision-language architecture, this model achieves state-of-the-art performance while maintaining exceptional efficiency
|
||||
with only 22B active parameters. Supporting 2M thinking tokens for advanced reasoning, the model represents a significant advancement in democratizing AI through sustainable and efficient architectures.
|
||||
\end{abstract}
|
||||
@@ -78,7 +78,7 @@ The rapid advancement of artificial intelligence has created an unprecedented de
|
||||
\subsection{Key Innovations}
|
||||
\begin{itemize}
|
||||
\item \textbf{Efficient Architecture}: 22B active parameters from 235B total
|
||||
\item \textbf{Specialized Training}: Optimized for visual reasoning & analysis
|
||||
\item \textbf{Specialized Training}: Optimized for visual reasoning \& analysis
|
||||
\item \textbf{Extended Context}: 131K context window
|
||||
\item \textbf{Thinking Mode}: 2M thinking tokens
|
||||
|
||||
@@ -117,7 +117,7 @@ The model employs a sophisticated Mixture of Experts architecture that activates
|
||||
during inference while maintaining 235B total parameters for enhanced capability.
|
||||
|
||||
\subsubsection{Attention Mechanism}
|
||||
Specialized attention mechanisms optimized for visual reasoning & analysis.
|
||||
Specialized attention mechanisms optimized for visual reasoning \& analysis.
|
||||
|
||||
\subsubsection{Thinking Mode}
|
||||
Advanced reasoning through extended thinking tokens (up to 2M), enabling:
|
||||
@@ -171,7 +171,7 @@ Latency (First Token) & 180 ms \\
|
||||
The model was trained on a carefully curated dataset comprising:
|
||||
\begin{itemize}
|
||||
\item High-quality filtered web data (50TB)
|
||||
\item Domain-specific corpora for visual reasoning & analysis
|
||||
\item Domain-specific corpora for visual reasoning \& analysis
|
||||
\item Synthetic data generation for edge cases
|
||||
\item Human feedback through RLHF
|
||||
\end{itemize}
|
||||
@@ -187,11 +187,13 @@ The model was trained on a carefully curated dataset comprising:
|
||||
\section{Use Cases and Applications}
|
||||
|
||||
\subsection{Primary Applications}
|
||||
\begin{itemize}
|
||||
\item UI/UX design analysis
|
||||
\item Architecture and layout planning
|
||||
\item Visual question answering
|
||||
\item Design system generation
|
||||
\item Accessibility evaluation
|
||||
\end{itemize}
|
||||
|
||||
\subsection{Integration Examples}
|
||||
|
||||
@@ -212,7 +214,7 @@ analysis = processor.decode(outputs[0])
|
||||
|
||||
\subsection{Sustainability Metrics}
|
||||
\begin{itemize}
|
||||
\item \textbf{Carbon Footprint}: 0.35 kg CO₂e per million inferences
|
||||
\item \textbf{Carbon Footprint}: 0.35 kg CO\textsubscript{2}e per million inferences
|
||||
\item \textbf{Energy Usage}: 8.0 kWh per day (1000 users)
|
||||
\item \textbf{Efficiency Gain}: 90\% reduction vs comparable models
|
||||
\end{itemize}
|
||||
@@ -286,7 +288,7 @@ INT4 & 55 GB & A100 80GB \\
|
||||
\section{Conclusion}
|
||||
|
||||
\textbf{Zen-Designer-Thinking} represents a significant advancement in AI democratization,
|
||||
delivering exceptional performance for visual reasoning & analysis while maintaining
|
||||
delivering exceptional performance for visual reasoning \& analysis while maintaining
|
||||
unprecedented efficiency. Through innovative architecture design and careful optimization,
|
||||
the model achieves a balance between capability and sustainability that sets a new standard
|
||||
for responsible AI development.
|
||||
@@ -296,8 +298,19 @@ for responsible AI development.
|
||||
We thank the open-source community, our research partners, and the teams at Hanzo AI and
|
||||
Zoo Labs Foundation for their contributions to this work.
|
||||
|
||||
\bibliographystyle{plain}
|
||||
\bibliography{references}
|
||||
\begin{thebibliography}{99}
|
||||
\bibitem{vaswani2017attention} Vaswani, A. et al. (2017). Attention Is All You Need. NeurIPS 2017.
|
||||
\bibitem{brown2020language} Brown, T. et al. (2020). Language Models are Few-Shot Learners. NeurIPS 2020.
|
||||
\bibitem{ouyang2022training} Ouyang, L. et al. (2022). Training Language Models to Follow Instructions with Human Feedback. NeurIPS 2022.
|
||||
\bibitem{ho2020denoising} Ho, J., Jain, A. and Abbeel, P. (2020). Denoising Diffusion Probabilistic Models. NeurIPS 2020.
|
||||
\bibitem{rombach2022high} Rombach, R. et al. (2022). High-Resolution Image Synthesis with Latent Diffusion Models. CVPR 2022.
|
||||
\bibitem{radford2021learning} Radford, A. et al. (2021). Learning Transferable Visual Models From Natural Language Supervision. ICML 2021.
|
||||
\bibitem{shazeer2017outrageously} Shazeer, N. et al. (2017). Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer. ICLR 2017.
|
||||
\bibitem{fedus2022switch} Fedus, W. et al. (2022). Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity. JMLR 2022.
|
||||
\bibitem{wei2022chain} Wei, J. et al. (2022). Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. NeurIPS 2022.
|
||||
\bibitem{rafailov2023direct} Rafailov, R. et al. (2023). Direct Preference Optimization: Your Language Model is Secretly a Reward Model. NeurIPS 2023.
|
||||
\bibitem{touvron2023llama} Touvron, H. et al. (2023). LLaMA: Open and Efficient Foundation Language Models. arXiv:2302.13971.
|
||||
\end{thebibliography}
|
||||
|
||||
\appendix
|
||||
|
||||
|
||||
Binary file not shown.
+3
-3
@@ -5,7 +5,7 @@
|
||||
\usepackage{graphicx}
|
||||
\usepackage{hyperref}
|
||||
\usepackage{listings}
|
||||
\usepackage{color}
|
||||
\usepackage[dvipsnames]{xcolor}
|
||||
\usepackage{booktabs}
|
||||
\usepackage{float}
|
||||
\usepackage{geometry}
|
||||
@@ -39,7 +39,7 @@
|
||||
}
|
||||
|
||||
\author{
|
||||
Hanzo AI Research Team\thanks{research@hanzo.ai} \and
|
||||
Antje Worring \and Hanzo AI Research Team\thanks{research@hanzo.ai} \and
|
||||
Zoo Labs Foundation\thanks{foundation@zoo.ngo}
|
||||
}
|
||||
|
||||
@@ -273,7 +273,7 @@ Overall directorial vision & 4.1 \\
|
||||
\bottomrule
|
||||
\end{tabular}
|
||||
\caption{Professional Cinematographer Evaluation (N=20 evaluators, 100 storyboards)}
|
||||
\end{table>
|
||||
\end{table}
|
||||
|
||||
\subsection{Downstream Video Generation Quality}
|
||||
|
||||
|
||||
Binary file not shown.
@@ -17,7 +17,7 @@
|
||||
|
||||
\title{\textbf{DSO: Decentralized Training Infrastructure for Zen}\\
|
||||
\large Technical Report v2025.06}
|
||||
\author{Zen LM Research Team\\
|
||||
\author{Antje Worring, Zach Kelling \\ Zen LM Research Team\\
|
||||
\texttt{research@zenlm.org}}
|
||||
\date{June 2025}
|
||||
|
||||
|
||||
Binary file not shown.
@@ -2,6 +2,10 @@
|
||||
\usepackage[utf8]{inputenc}
|
||||
\usepackage[T1]{fontenc}
|
||||
\usepackage{amsmath,amsfonts,amssymb}
|
||||
\usepackage{amsthm}
|
||||
\newtheorem{theorem}{Theorem}
|
||||
\newtheorem{lemma}[theorem]{Lemma}
|
||||
\newtheorem{definition}[theorem]{Definition}
|
||||
\usepackage{graphicx}
|
||||
\usepackage{hyperref}
|
||||
\usepackage{listings}
|
||||
@@ -17,7 +21,7 @@
|
||||
|
||||
\title{\textbf{DSO: Decentralized Semantic Optimization Protocol}\\
|
||||
\large Technical Report v2025.06}
|
||||
\author{Zen LM Research Team\\
|
||||
\author{Antje Worring, Zach Kelling \\ Zen LM Research Team\\
|
||||
\texttt{research@zenlm.org}}
|
||||
\date{June 2025}
|
||||
|
||||
|
||||
Binary file not shown.
@@ -16,7 +16,7 @@
|
||||
\title{\textbf{Zen-Dub-Live: Real-Time Streaming AI Dubbing\\
|
||||
for Live Video Content}\\[0.5em]
|
||||
\large Technical Whitepaper v2025.04}
|
||||
\author{Zen LM Research Team\\
|
||||
\author{Antje Worring, Zach Kelling \\ Zen LM Research Team\\
|
||||
\texttt{research@zenlm.org}\\
|
||||
\href{https://papers.zenlm.org}{papers.zenlm.org}}
|
||||
\date{April 2025}
|
||||
@@ -122,7 +122,7 @@ Speaker identity is preserved without pre-enrollment using a zero-shot voice clo
|
||||
|
||||
\begin{equation}
|
||||
v_t = E_\phi\left(a_{t-3s:t}\right)
|
||||
\end{equation>
|
||||
\end{equation}
|
||||
|
||||
The voice embedding conditions the vocoder through cross-attention in the synthesis layers, transferring spectral envelope characteristics (timbre, formant patterns) while allowing the translated phoneme sequence to drive articulation. Speaker similarity is measured by cosine distance between voice embeddings of original and dubbed speech.
|
||||
|
||||
|
||||
Binary file not shown.
@@ -15,7 +15,7 @@
|
||||
|
||||
\title{\textbf{Zen-Dub: AI Voice Dubbing and Localization}\\
|
||||
\large Technical Whitepaper v2025.03}
|
||||
\author{Zen LM Research Team\\
|
||||
\author{Antje Worring, Zach Kelling \\ Zen LM Research Team\\
|
||||
\texttt{research@zenlm.org}\\
|
||||
\href{https://papers.zenlm.org}{papers.zenlm.org}}
|
||||
\date{March 2025}
|
||||
|
||||
@@ -1,323 +0,0 @@
|
||||
\documentclass[11pt,a4paper]{article}
|
||||
\usepackage[utf8]{inputenc}
|
||||
\usepackage[T1]{fontenc}
|
||||
\usepackage{amsmath,amsfonts,amssymb}
|
||||
\usepackage{graphicx}
|
||||
\usepackage{hyperref}
|
||||
\usepackage{listings}
|
||||
\usepackage{color}
|
||||
\usepackage{booktabs}
|
||||
\usepackage{float}
|
||||
\usepackage{geometry}
|
||||
\geometry{margin=1in}
|
||||
|
||||
% Color definitions
|
||||
\definecolor{zenblue}{RGB}{41,121,255}
|
||||
\definecolor{zengreen}{RGB}{52,199,89}
|
||||
\definecolor{zenorange}{RGB}{255,149,0}
|
||||
\definecolor{codegray}{RGB}{245,245,245}
|
||||
|
||||
% Hyperref setup
|
||||
\hypersetup{
|
||||
colorlinks=true,
|
||||
linkcolor=zenblue,
|
||||
urlcolor=zenblue,
|
||||
citecolor=zenblue
|
||||
}
|
||||
|
||||
% Code listing setup
|
||||
\lstset{
|
||||
backgroundcolor=\color{codegray},
|
||||
basicstyle=\ttfamily\small,
|
||||
breaklines=true,
|
||||
captionpos=b,
|
||||
frame=single,
|
||||
numbers=left,
|
||||
numberstyle=\tiny\color{gray}
|
||||
}
|
||||
|
||||
\title{
|
||||
\vspace{-2cm}
|
||||
\Large \textbf{Zen AI Model Family} \\
|
||||
\vspace{0.5cm}
|
||||
\Huge \textbf{Zen-Eco} \\
|
||||
\vspace{0.3cm}
|
||||
\large Consumer Hardware \\
|
||||
\vspace{0.5cm}
|
||||
\normalsize Technical Whitepaper v1.0
|
||||
}
|
||||
|
||||
\author{
|
||||
Zach Kelling\thanks{zach@lux.network} \\
|
||||
\texttt{research@hanzo.ai} \\
|
||||
\\
|
||||
Zoo Labs Foundation \\
|
||||
\texttt{foundation@zoolabs.org}
|
||||
}
|
||||
|
||||
\date{September 2025}
|
||||
|
||||
\begin{document}
|
||||
|
||||
\maketitle
|
||||
|
||||
\begin{abstract}
|
||||
We present \textbf{Zen-Eco}, a 4B parameter model optimized for consumer hardware.
|
||||
Built upon zen-3B, this model achieves state-of-the-art performance while maintaining exceptional efficiency
|
||||
with only 4B active parameters. Supporting 128K thinking tokens for advanced reasoning, the model represents a significant advancement in democratizing AI through sustainable and efficient architectures.
|
||||
\end{abstract}
|
||||
|
||||
\tableofcontents
|
||||
\newpage
|
||||
|
||||
\section{Introduction}
|
||||
|
||||
The rapid advancement of artificial intelligence has created an unprecedented demand for models that balance capability with efficiency.
|
||||
\textbf{Zen-Eco} addresses this challenge by delivering enterprise-grade performance while maintaining a minimal computational footprint.
|
||||
|
||||
\subsection{Key Innovations}
|
||||
\begin{itemize}
|
||||
\item \textbf{Efficient Architecture}: 4B active parameters from 4B total
|
||||
\item \textbf{Specialized Training}: Optimized for consumer hardware
|
||||
\item \textbf{Extended Context}: 32K context window
|
||||
\item \textbf{Thinking Mode}: 128K thinking tokens
|
||||
|
||||
|
||||
\end{itemize}
|
||||
|
||||
\section{Architecture}
|
||||
|
||||
\subsection{Model Design}
|
||||
|
||||
Zen-Eco is based on the zen-3B architecture with several key modifications:
|
||||
|
||||
\begin{table}[H]
|
||||
\centering
|
||||
\begin{tabular}{ll}
|
||||
\toprule
|
||||
\textbf{Component} & \textbf{Specification} \\
|
||||
\midrule
|
||||
Total Parameters & 4B \\
|
||||
Active Parameters & 4B \\
|
||||
Base Model & zen-3B \\
|
||||
Context Length & 32K \\
|
||||
Thinking Tokens & 128K \\
|
||||
|
||||
|
||||
Architecture Type & Transformer \\
|
||||
\bottomrule
|
||||
\end{tabular}
|
||||
\caption{Zen-Eco Architecture Specifications}
|
||||
\end{table}
|
||||
|
||||
\subsection{Technical Innovations}
|
||||
|
||||
\subsubsection{Mixture of Experts (MoE)}
|
||||
The model uses a dense architecture with all parameters active during inference, optimized for maximum performance per parameter.
|
||||
|
||||
\subsubsection{Attention Mechanism}
|
||||
Extended attention mechanisms support up to 32K context length with efficient KV-cache management.
|
||||
|
||||
\subsubsection{Thinking Mode}
|
||||
Advanced reasoning through extended thinking tokens (up to 128K), enabling:
|
||||
\begin{itemize}
|
||||
\item Step-by-step problem decomposition
|
||||
\item Self-correction and verification
|
||||
\item Complex multi-step reasoning
|
||||
\item Internal deliberation before response
|
||||
\end{itemize}
|
||||
|
||||
\section{Performance Benchmarks}
|
||||
|
||||
\subsection{Evaluation Results}
|
||||
|
||||
|
||||
\begin{table}[H]
|
||||
\centering
|
||||
\begin{tabular}{lc}
|
||||
\toprule
|
||||
\textbf{Benchmark} & \textbf{Score} \\
|
||||
\midrule
|
||||
MMLU & 62.3\% \\
|
||||
HumanEval & 35.2\% \\
|
||||
GSM8K & 74.8\% \\
|
||||
HellaSwag & 71.6\% \\
|
||||
\bottomrule
|
||||
\end{tabular}
|
||||
\caption{Language Understanding Benchmarks}
|
||||
\end{table}
|
||||
|
||||
\subsection{Efficiency Metrics}
|
||||
|
||||
\begin{table}[H]
|
||||
\centering
|
||||
\begin{tabular}{ll}
|
||||
\toprule
|
||||
\textbf{Metric} & \textbf{Value} \\
|
||||
\midrule
|
||||
Inference Speed & 250 tokens/sec \\
|
||||
Memory Usage (INT4) & 8 GB \\
|
||||
Energy Efficiency & 95\% reduction \\
|
||||
Latency (First Token) & 35 ms \\
|
||||
\bottomrule
|
||||
\end{tabular}
|
||||
\caption{Efficiency Metrics}
|
||||
\end{table}
|
||||
|
||||
\section{Training Methodology}
|
||||
|
||||
\subsection{Dataset}
|
||||
The model was trained on a carefully curated dataset comprising:
|
||||
\begin{itemize}
|
||||
\item High-quality filtered web data (2TB)
|
||||
\item Domain-specific corpora for consumer hardware
|
||||
\item Synthetic data generation for edge cases
|
||||
\item Human feedback through RLHF
|
||||
\end{itemize}
|
||||
|
||||
\subsection{Training Process}
|
||||
\begin{enumerate}
|
||||
\item \textbf{Pretraining}: 2 trillion tokens over 14 days on 8x A100
|
||||
\item \textbf{Supervised Fine-tuning}: Task-specific optimization
|
||||
\item \textbf{RLHF}: Alignment with human preferences
|
||||
\item \textbf{Constitutional AI}: Safety and helpfulness optimization
|
||||
\end{enumerate}
|
||||
|
||||
\section{Use Cases and Applications}
|
||||
|
||||
\subsection{Primary Applications}
|
||||
\item Conversational AI and chatbots
|
||||
\item Content generation and summarization
|
||||
\item Code completion and review
|
||||
\item Educational assistance
|
||||
\item Research and analysis
|
||||
|
||||
\subsection{Integration Examples}
|
||||
|
||||
\begin{lstlisting}[language=Python, caption=Basic Usage Example]
|
||||
from transformers import AutoModelForCausalLM, AutoTokenizer
|
||||
|
||||
# Load model and tokenizer
|
||||
model = AutoModelForCausalLM.from_pretrained("zenlm/zen-eco-4b-instruct")
|
||||
tokenizer = AutoTokenizer.from_pretrained("zenlm/zen-eco-4b-instruct")
|
||||
|
||||
# Generate response
|
||||
inputs = tokenizer("Explain quantum computing", return_tensors="pt")
|
||||
outputs = model.generate(**inputs, max_length=100)
|
||||
response = tokenizer.decode(outputs[0])
|
||||
\end{lstlisting}
|
||||
|
||||
\section{Environmental Impact}
|
||||
|
||||
\subsection{Sustainability Metrics}
|
||||
\begin{itemize}
|
||||
\item \textbf{Carbon Footprint}: 0.05 kg CO₂e per million inferences
|
||||
\item \textbf{Energy Usage}: 1.2 kWh per day (1000 users)
|
||||
\item \textbf{Efficiency Gain}: 95\% reduction vs comparable models
|
||||
\end{itemize}
|
||||
|
||||
\subsection{Green AI Commitment}
|
||||
Zen AI models are designed with sustainability as a core principle, achieving industry-leading efficiency
|
||||
through architectural innovations and optimization techniques.
|
||||
|
||||
\section{Safety and Alignment}
|
||||
|
||||
\subsection{Safety Measures}
|
||||
\begin{itemize}
|
||||
\item Constitutional AI training for harmlessness
|
||||
\item Comprehensive red-teaming and adversarial testing
|
||||
\item Built-in safety filters and guardrails
|
||||
\item Regular safety audits and updates
|
||||
\end{itemize}
|
||||
|
||||
\subsection{Ethical Considerations}
|
||||
The model has been developed with careful attention to:
|
||||
\begin{itemize}
|
||||
\item Bias mitigation through diverse training data
|
||||
\item Transparency in capabilities and limitations
|
||||
\item Privacy-preserving deployment options
|
||||
\item Responsible AI principles alignment
|
||||
\end{itemize}
|
||||
|
||||
\section{Deployment Options}
|
||||
|
||||
\subsection{Available Formats}
|
||||
\begin{itemize}
|
||||
\item \textbf{SafeTensors}: Original precision weights
|
||||
\item \textbf{GGUF}: Quantized formats (Q4\_K\_M, Q5\_K\_M, Q8\_0)
|
||||
\item \textbf{MLX}: Apple Silicon optimization (4-bit, 8-bit)
|
||||
\item \textbf{ONNX}: Cross-platform deployment (coming soon)
|
||||
\end{itemize}
|
||||
|
||||
\subsection{Hardware Requirements}
|
||||
\begin{table}[H]
|
||||
\centering
|
||||
\begin{tabular}{lll}
|
||||
\toprule
|
||||
\textbf{Precision} & \textbf{Memory} & \textbf{Recommended Hardware} \\
|
||||
\midrule
|
||||
FP16 & 8 GB & RTX 3070 \\
|
||||
INT8 & 4 GB & RTX 3060 \\
|
||||
INT4 & 8 GB & M2 MacBook Air \\
|
||||
\bottomrule
|
||||
\end{tabular}
|
||||
\caption{Hardware Requirements by Precision}
|
||||
\end{table}
|
||||
|
||||
\section{Future Work}
|
||||
|
||||
\subsection{Planned Improvements}
|
||||
\begin{itemize}
|
||||
\item Extended context windows (up to 1M tokens)
|
||||
\item Enhanced multimodal capabilities
|
||||
\item Improved efficiency through further optimization
|
||||
\item Expanded language support
|
||||
\end{itemize}
|
||||
|
||||
\subsection{Research Directions}
|
||||
\begin{itemize}
|
||||
\item Advanced reasoning mechanisms
|
||||
\item Self-supervised learning improvements
|
||||
\item Zero-shot generalization enhancement
|
||||
\item Continual learning capabilities
|
||||
\end{itemize}
|
||||
|
||||
\section{Conclusion}
|
||||
|
||||
\textbf{Zen-Eco} represents a significant advancement in AI democratization,
|
||||
delivering exceptional performance for consumer hardware while maintaining
|
||||
unprecedented efficiency. Through innovative architecture design and careful optimization,
|
||||
the model achieves a balance between capability and sustainability that sets a new standard
|
||||
for responsible AI development.
|
||||
|
||||
\section*{Acknowledgments}
|
||||
|
||||
We thank the open-source community, our research partners, and the teams at Hanzo AI and
|
||||
Zoo Labs Foundation for their contributions to this work.
|
||||
|
||||
\bibliographystyle{plain}
|
||||
\bibliography{references}
|
||||
|
||||
\appendix
|
||||
|
||||
\section{Model Card}
|
||||
|
||||
\begin{table}[H]
|
||||
\centering
|
||||
\begin{tabular}{ll}
|
||||
\toprule
|
||||
\textbf{Field} & \textbf{Value} \\
|
||||
\midrule
|
||||
Model Name & Zen-Eco \\
|
||||
Version & 1.0.0 \\
|
||||
Release Date & September 2025 \\
|
||||
License & Apache 2.0 \\
|
||||
Repository & \href{https://huggingface.co/zenlm/zen-eco-4b-instruct}{huggingface.co/zenlm/zen-eco-4b-instruct} \\
|
||||
Documentation & \href{https://github.com/zenlm/zen}{github.com/zenlm/zen} \\
|
||||
Contact & research@hanzo.ai \\
|
||||
\bottomrule
|
||||
\end{tabular}
|
||||
\caption{Model Card Information}
|
||||
\end{table}
|
||||
|
||||
\end{document}
|
||||
Binary file not shown.
@@ -15,7 +15,7 @@
|
||||
|
||||
\title{\textbf{Zen Embeddings: Dense Retrieval and Semantic Search}\\
|
||||
\large Technical Report v2025.06}
|
||||
\author{Zen LM Research Team\\
|
||||
\author{Antje Worring, Zach Kelling \\ Zen LM Research Team\\
|
||||
\texttt{research@zenlm.org}}
|
||||
\date{June 2025}
|
||||
|
||||
|
||||
Binary file not shown.
@@ -17,7 +17,7 @@
|
||||
|
||||
\title{\textbf{Zen Enterprise: Deployment Architecture and Operations}\\
|
||||
\large Technical Report v2025.08}
|
||||
\author{Zen LM Research Team\\
|
||||
\author{Antje Worring, Zach Kelling \\ Zen LM Research Team\\
|
||||
\texttt{research@zenlm.org}}
|
||||
\date{August 2025}
|
||||
|
||||
@@ -203,7 +203,7 @@ Enterprise & 99.99\% & 200ms & 8s & Critical \\
|
||||
Dedicated & 99.99\% & 150ms & 5s & Reserved capacity \\
|
||||
\bottomrule
|
||||
\end{tabular}
|
||||
\end{table>
|
||||
\end{table}
|
||||
|
||||
\subsection{SLA Monitoring and Alerting}
|
||||
|
||||
|
||||
Binary file not shown.
@@ -15,7 +15,7 @@
|
||||
|
||||
\title{\textbf{Zen Financial: AI for Capital Markets and Finance}\\
|
||||
\large Technical Report v2025.09}
|
||||
\author{Zen LM Research Team\\
|
||||
\author{Antje Worring, Zach Kelling \\ Zen LM Research Team\\
|
||||
\texttt{research@zenlm.org}}
|
||||
\date{September 2025}
|
||||
|
||||
|
||||
Binary file not shown.
+1
-1
@@ -17,7 +17,7 @@
|
||||
|
||||
\title{\textbf{Fine-Tuning Zen Models: Methods and Best Practices}\\
|
||||
\large Technical Report v2025.04}
|
||||
\author{Zen LM Research Team\\
|
||||
\author{Antje Worring, Zach Kelling \\ Zen LM Research Team\\
|
||||
\texttt{research@zenlm.org}}
|
||||
\date{April 2025}
|
||||
|
||||
|
||||
Binary file not shown.
+5
-5
@@ -5,7 +5,7 @@
|
||||
\usepackage{graphicx}
|
||||
\usepackage{hyperref}
|
||||
\usepackage{listings}
|
||||
\usepackage{color}
|
||||
\usepackage[dvipsnames]{xcolor}
|
||||
\usepackage{booktabs}
|
||||
\usepackage{float}
|
||||
\usepackage{geometry}
|
||||
@@ -39,7 +39,7 @@
|
||||
}
|
||||
|
||||
\author{
|
||||
Hanzo AI Research Team\thanks{research@hanzo.ai} \and
|
||||
Antje Worring \and Hanzo AI Research Team\thanks{research@hanzo.ai} \and
|
||||
Zoo Labs Foundation\thanks{foundation@zoo.ngo}
|
||||
}
|
||||
|
||||
@@ -263,7 +263,7 @@ V2A-Mapper & 4.0 & 3.9 & 4.0 \\
|
||||
\bottomrule
|
||||
\end{tabular}
|
||||
\caption{Mean Opinion Score Evaluation (1--5 scale, N=15 audio engineers)}
|
||||
\end{table>
|
||||
\end{table}
|
||||
|
||||
\subsection{AVSync Benchmark (Synchronization)}
|
||||
|
||||
@@ -283,7 +283,7 @@ TempoFoley & 68 & 89.4 \\
|
||||
\bottomrule
|
||||
\end{tabular}
|
||||
\caption{AVSync Synchronization Benchmark}
|
||||
\end{table>
|
||||
\end{table}
|
||||
|
||||
\subsection{Spatial Audio Quality}
|
||||
|
||||
@@ -303,7 +303,7 @@ Distance rank correlation & 0.84 & 0.91 \\
|
||||
\bottomrule
|
||||
\end{tabular}
|
||||
\caption{Spatial Audio Localization Accuracy}
|
||||
\end{table>
|
||||
\end{table}
|
||||
|
||||
\section{Applications}
|
||||
|
||||
|
||||
Binary file not shown.
@@ -16,7 +16,7 @@
|
||||
\title{\textbf{Zen-Guard-Gen: Generative Safety Classification\\
|
||||
with Natural Language Explanations and Policy References}\\[0.5em]
|
||||
\large Technical Whitepaper v2025.05}
|
||||
\author{Zen LM Research Team\\
|
||||
\author{Antje Worring, Zach Kelling \\ Zen LM Research Team\\
|
||||
\texttt{research@zenlm.org}\\
|
||||
\href{https://papers.zenlm.org}{papers.zenlm.org}}
|
||||
\date{May 2025}
|
||||
|
||||
Binary file not shown.
@@ -16,7 +16,7 @@
|
||||
\title{\textbf{Zen-Guard-Stream: Token-Level Streaming Safety Filtering\\
|
||||
for Real-Time Generative AI Systems}\\[0.5em]
|
||||
\large Technical Whitepaper v2025.05}
|
||||
\author{Zen LM Research Team\\
|
||||
\author{Antje Worring, Zach Kelling \\ Zen LM Research Team\\
|
||||
\texttt{research@zenlm.org}\\
|
||||
\href{https://papers.zenlm.org}{papers.zenlm.org}}
|
||||
\date{May 2025}
|
||||
|
||||
@@ -1,143 +0,0 @@
|
||||
\documentclass[11pt]{article}
|
||||
\usepackage[margin=1in]{geometry}
|
||||
\usepackage{graphicx}
|
||||
\usepackage{hyperref}
|
||||
\usepackage{amsmath}
|
||||
\usepackage{booktabs}
|
||||
|
||||
\title{Zen-Guard: Multilingual Safety Moderation for AI Systems\\
|
||||
\large Technical Whitepaper}
|
||||
\author{Zach Kelling\thanks{zach@lux.network} \and Hanzo Industries \and Lux Industries \and Zoo Labs Foundation}
|
||||
\date{September 2025}
|
||||
|
||||
\begin{document}
|
||||
|
||||
\maketitle
|
||||
|
||||
\begin{abstract}
|
||||
Zen-Guard represents a comprehensive safety moderation solution for AI systems, offering both generative and streaming variants for real-time content filtering. Built upon advanced architectures with support for 119 languages, Zen-Guard provides three-tier severity classification across 9 safety categories. The models achieve 96.8\% accuracy with minimal false positives, enabling robust content moderation at scale.
|
||||
\end{abstract}
|
||||
|
||||
\section{Introduction}
|
||||
|
||||
As AI systems become increasingly prevalent, ensuring safe and appropriate content generation is paramount. Zen-Guard addresses this challenge through specialized models optimized for different deployment scenarios:
|
||||
|
||||
\begin{itemize}
|
||||
\item \textbf{Zen-Guard-Gen (8B)}: Generative safety classification
|
||||
\item \textbf{Zen-Guard-Stream (4B)}: Real-time token-level monitoring
|
||||
\end{itemize}
|
||||
|
||||
\section{Architecture}
|
||||
|
||||
\subsection{Model Variants}
|
||||
|
||||
\begin{table}[h]
|
||||
\centering
|
||||
\begin{tabular}{lcccc}
|
||||
\toprule
|
||||
Model & Parameters & Type & Languages & Latency \\
|
||||
\midrule
|
||||
Guard-Gen-8B & 8B & Generative & 119 & 120ms \\
|
||||
Guard-Stream-4B & 4B & Streaming & 119 & 5ms/token \\
|
||||
\bottomrule
|
||||
\end{tabular}
|
||||
\caption{Zen-Guard model specifications}
|
||||
\end{table}
|
||||
|
||||
\subsection{Safety Categories}
|
||||
|
||||
The models classify content across 9 primary categories:
|
||||
\begin{enumerate}
|
||||
\item Violent content and instructions
|
||||
\item Non-violent illegal activities
|
||||
\item Sexual content or acts
|
||||
\item Personally identifiable information
|
||||
\item Suicide and self-harm
|
||||
\item Unethical acts and discrimination
|
||||
\item Politically sensitive topics
|
||||
\item Copyright violations
|
||||
\item Jailbreak attempts
|
||||
\end{enumerate}
|
||||
|
||||
\section{Performance Metrics}
|
||||
|
||||
\subsection{Benchmark Results}
|
||||
|
||||
\begin{table}[h]
|
||||
\centering
|
||||
\begin{tabular}{lcccc}
|
||||
\toprule
|
||||
Metric & Guard-Gen & Guard-Stream & Industry Avg \\
|
||||
\midrule
|
||||
Accuracy & 96.8\% & 95.2\% & 92.1\% \\
|
||||
F1 Score & 94.2\% & 93.1\% & 89.5\% \\
|
||||
False Positive & 2.1\% & 2.8\% & 5.3\% \\
|
||||
Latency & 120ms & 5ms & 200ms \\
|
||||
\bottomrule
|
||||
\end{tabular}
|
||||
\caption{Performance comparison}
|
||||
\end{table}
|
||||
|
||||
\subsection{Multilingual Performance}
|
||||
|
||||
Zen-Guard maintains consistent performance across all 119 supported languages:
|
||||
\begin{itemize}
|
||||
\item English: 97.2\% accuracy
|
||||
\item Chinese: 96.5\% accuracy
|
||||
\item Spanish: 96.1\% accuracy
|
||||
\item Other languages: 95.8\% average
|
||||
\end{itemize}
|
||||
|
||||
\section{Deployment}
|
||||
|
||||
\subsection{Integration Options}
|
||||
|
||||
\begin{enumerate}
|
||||
\item \textbf{API Integration}: REST/GraphQL endpoints
|
||||
\item \textbf{Edge Deployment}: Optimized for local inference
|
||||
\item \textbf{Streaming Integration}: Real-time token filtering
|
||||
\item \textbf{Batch Processing}: High-throughput moderation
|
||||
\end{enumerate}
|
||||
|
||||
\subsection{Resource Requirements}
|
||||
|
||||
\begin{itemize}
|
||||
\item Guard-Gen-8B: 16GB VRAM (FP16), 8GB (INT8)
|
||||
\item Guard-Stream-4B: 8GB VRAM (FP16), 4GB (INT8)
|
||||
\item CPU: 8+ cores recommended
|
||||
\item Throughput: 1000+ requests/second
|
||||
\end{itemize}
|
||||
|
||||
\section{Use Cases}
|
||||
|
||||
\subsection{Application Scenarios}
|
||||
|
||||
\begin{itemize}
|
||||
\item \textbf{Chat Applications}: Real-time message filtering
|
||||
\item \textbf{Content Platforms}: User-generated content moderation
|
||||
\item \textbf{Educational Systems}: Safe learning environments
|
||||
\item \textbf{Enterprise AI}: Compliance and safety assurance
|
||||
\item \textbf{Gaming}: Community interaction monitoring
|
||||
\end{itemize}
|
||||
|
||||
\section{Environmental Impact}
|
||||
|
||||
\begin{itemize}
|
||||
\item Energy Usage: 92\% less than comparable models
|
||||
\item Carbon Footprint: 0.8kg CO₂/month per instance
|
||||
\item Optimization: INT8 quantization reduces energy by 50\%
|
||||
\end{itemize}
|
||||
|
||||
\section{Conclusion}
|
||||
|
||||
Zen-Guard provides comprehensive, multilingual safety moderation with industry-leading performance. The dual-model approach ensures flexibility for both batch and real-time applications while maintaining high accuracy and low false positive rates.
|
||||
|
||||
\section{References}
|
||||
|
||||
\begin{enumerate}
|
||||
\item Zen-Guard Architecture Technical Report (2025)
|
||||
\item Multilingual Safety Moderation Benchmarks
|
||||
\item Real-time Content Filtering Systems
|
||||
\end{enumerate}
|
||||
|
||||
\end{document}
|
||||
Binary file not shown.
@@ -16,7 +16,7 @@
|
||||
\title{\textbf{Reducing Hallucinations in Zen Models: Retrieval Augmentation,\\
|
||||
Confidence Calibration, and Citation-Grounded Generation}\\
|
||||
\large Technical Report v2025.08}
|
||||
\author{Zen LM Research Team\\
|
||||
\author{Antje Worring, Zach Kelling \\ Zen LM Research Team\\
|
||||
\texttt{research@zenlm.org}}
|
||||
\date{August 2025}
|
||||
|
||||
@@ -342,7 +342,7 @@ Mathematics & 9.6 & 5.4 & $-43.8\%$ \\
|
||||
\bottomrule
|
||||
\end{tabular}
|
||||
\label{tab:domain_results}
|
||||
\end{table>
|
||||
\end{table}
|
||||
|
||||
\subsection{Citation Accuracy}
|
||||
|
||||
|
||||
Binary file not shown.
@@ -26,7 +26,7 @@
|
||||
\title{\textbf{Zen Hardware Optimization: GPU, TPU, and Edge Deployment\\
|
||||
FlashAttention, Custom CUDA Kernels, and Architecture-Aware Quantization}\\
|
||||
\large Technical Report v2025.06}
|
||||
\author{Zen LM Research Team\\
|
||||
\author{Antje Worring, Zach Kelling \\ Zen LM Research Team\\
|
||||
\texttt{research@zenlm.org}}
|
||||
\date{June 2025}
|
||||
|
||||
|
||||
Binary file not shown.
@@ -15,7 +15,7 @@
|
||||
|
||||
\title{\textbf{Zen Inference Optimization: Serving at Scale}\\
|
||||
\large Technical Report v2025.05}
|
||||
\author{Zen LM Research Team\\
|
||||
\author{Antje Worring, Zach Kelling \\ Zen LM Research Team\\
|
||||
\texttt{research@zenlm.org}}
|
||||
\date{May 2025}
|
||||
|
||||
|
||||
Binary file not shown.
@@ -16,7 +16,7 @@
|
||||
\title{\textbf{Knowledge Distillation for Efficient Zen Models:\\
|
||||
Semantic-Preserving Multi-Teacher Distillation}\\
|
||||
\large Technical Report v2025.07}
|
||||
\author{Zen LM Research Team\\
|
||||
\author{Antje Worring, Zach Kelling \\ Zen LM Research Team\\
|
||||
\texttt{research@zenlm.org}}
|
||||
\date{July 2025}
|
||||
|
||||
|
||||
Binary file not shown.
+1
-1
@@ -15,7 +15,7 @@
|
||||
|
||||
\title{\textbf{Zen Legal: AI-Assisted Legal Research and Analysis}\\
|
||||
\large Technical Report v2025.10}
|
||||
\author{Zen LM Research Team\\
|
||||
\author{Antje Worring, Zach Kelling \\ Zen LM Research Team\\
|
||||
\texttt{research@zenlm.org}}
|
||||
\date{October 2025}
|
||||
|
||||
|
||||
Binary file not shown.
@@ -16,7 +16,7 @@
|
||||
\title{\textbf{Zen-Live: A Real-Time Conversational AI Model\\
|
||||
with Native Turn Management and Voice Interaction}\\[0.5em]
|
||||
\large Technical Whitepaper v2025.03}
|
||||
\author{Zen LM Research Team\\
|
||||
\author{Antje Worring, Zach Kelling \\ Zen LM Research Team\\
|
||||
\texttt{research@zenlm.org}\\
|
||||
\href{https://papers.zenlm.org}{papers.zenlm.org}}
|
||||
\date{March 2025}
|
||||
|
||||
Binary file not shown.
@@ -17,7 +17,7 @@
|
||||
|
||||
\title{\textbf{Mathematical Reasoning in Zen Models}\\
|
||||
\large Technical Report v2025.08}
|
||||
\author{Zen LM Research Team\\
|
||||
\author{Antje Worring, Zach Kelling \\ Zen LM Research Team\\
|
||||
\texttt{research@zenlm.org}}
|
||||
\date{August 2025}
|
||||
|
||||
@@ -304,7 +304,7 @@ Total & 200 & \\
|
||||
\bottomrule
|
||||
\end{tabular}
|
||||
\caption{Error taxonomy for Zen-7B on MATH Level 5. PRM detects most computational errors.}
|
||||
\end{table>
|
||||
\end{table}
|
||||
|
||||
The PRM is most effective at detecting arithmetic and algebraic errors (step produces
|
||||
numerically inconsistent result) and least effective at detecting missing cases or
|
||||
|
||||
Binary file not shown.
@@ -16,7 +16,7 @@
|
||||
|
||||
\title{\textbf{Zen-Max: Maximum Capability via Mixture-of-Experts Architecture}\\
|
||||
\large Technical Report v2025.03}
|
||||
\author{Zen LM Research Team\\
|
||||
\author{Antje Worring, Zach Kelling \\ Zen LM Research Team\\
|
||||
\texttt{research@zenlm.org}}
|
||||
\date{March 2025}
|
||||
|
||||
|
||||
@@ -1,287 +0,0 @@
|
||||
\documentclass[11pt,a4paper]{article}
|
||||
\usepackage[utf8]{inputenc}
|
||||
\usepackage[T1]{fontenc}
|
||||
\usepackage{amsmath,amsfonts,amssymb}
|
||||
\usepackage{graphicx}
|
||||
\usepackage{hyperref}
|
||||
\usepackage{listings}
|
||||
\usepackage{color}
|
||||
\usepackage{booktabs}
|
||||
\usepackage{float}
|
||||
\usepackage{geometry}
|
||||
\geometry{margin=1in}
|
||||
\definecolor{zenblue}{RGB}{41,121,255}
|
||||
\hypersetup{colorlinks=true,linkcolor=zenblue,urlcolor=zenblue,citecolor=zenblue}
|
||||
|
||||
\title{\textbf{Zen Medical: Clinical AI and Healthcare Applications}\\
|
||||
\large Technical Report v2025.09}
|
||||
\author{Zen LM Research Team\\
|
||||
\texttt{research@zenlm.org}}
|
||||
\date{September 2025}
|
||||
|
||||
\begin{document}
|
||||
\maketitle
|
||||
|
||||
\begin{abstract}
|
||||
We present Zen Medical, a suite of clinical AI models and deployment infrastructure built on the Zen MoDE (Mixture of Distilled Experts) backbone. Zen Medical addresses the unique requirements of healthcare AI: evidence-grounded generation, HIPAA-compliant on-premise deployment, clinical reasoning chain transparency, and integration with electronic health records (EHR) systems. On standardized benchmarks, Zen Medical achieves: MedQA (USMLE Step 1--3) 89.3\%, USMLE four-step 87.2\%, PubMedQA 76.4\%, and clinical NER F1 0.918. We describe the medical fine-tuning methodology, safety guardrails, deployment architecture, and clinical validation results.
|
||||
\end{abstract}
|
||||
|
||||
\section{Introduction}
|
||||
|
||||
Healthcare AI presents challenges qualitatively different from general-purpose AI deployment. Errors have direct patient safety consequences. Regulatory requirements (HIPAA, FDA, EU MDR) constrain data handling and system behavior. Clinical workflows demand seamless EHR integration. Clinicians require transparent reasoning, not opaque outputs.
|
||||
|
||||
Zen Medical addresses these requirements through a ground-up approach to medical AI:
|
||||
|
||||
\begin{enumerate}
|
||||
\item \textbf{Medical knowledge fine-tuning}: Specialized training on curated clinical literature, guidelines, and case databases.
|
||||
\item \textbf{Evidence citation}: Every clinical claim cites a source—PubMed article, clinical guideline, or drug label.
|
||||
\item \textbf{Clinical reasoning chains}: Step-by-step diagnostic and therapeutic reasoning exposed to the user.
|
||||
\item \textbf{HIPAA-compliant deployment}: On-premise Kubernetes deployment with end-to-end encryption, audit logging, and zero data egress.
|
||||
\item \textbf{Safety guardrails}: Conservative uncertainty quantification, mandatory escalation triggers, and clinician-in-the-loop design.
|
||||
\end{enumerate}
|
||||
|
||||
\section{Medical Knowledge Fine-tuning}
|
||||
|
||||
\subsection{Training Corpus}
|
||||
|
||||
Zen Medical is fine-tuned on a curated 680-billion-token medical corpus assembled from:
|
||||
|
||||
\begin{table}[H]
|
||||
\centering
|
||||
\caption{Medical training corpus composition}
|
||||
\label{tab:corpus}
|
||||
\begin{tabular}{lcc}
|
||||
\toprule
|
||||
Source & Tokens (B) & Weight \\
|
||||
\midrule
|
||||
PubMed full-text articles & 142 & 0.21 \\
|
||||
Clinical guidelines (AHA, ACS, WHO, etc.) & 8.4 & 0.08 \\
|
||||
Medical textbooks (licensed) & 24 & 0.12 \\
|
||||
De-identified EHR notes & 184 & 0.27 \\
|
||||
Drug labels (FDA, EMA) & 6.8 & 0.05 \\
|
||||
Medical QA databases & 38 & 0.10 \\
|
||||
Imaging reports (radiology, pathology) & 92 & 0.13 \\
|
||||
Synthetic clinical dialogues & 18 & 0.04 \\
|
||||
\bottomrule
|
||||
\end{tabular}
|
||||
\end{table}
|
||||
|
||||
All EHR data was processed through a multi-layer de-identification pipeline achieving $<$0.001\% residual PHI (Protected Health Information) retention rate, verified by clinical privacy auditors.
|
||||
|
||||
\subsection{Instruction Fine-tuning}
|
||||
|
||||
Following corpus pre-training, Zen Medical undergoes instruction fine-tuning on 4.2 million medical instruction-response pairs covering:
|
||||
|
||||
\begin{itemize}
|
||||
\item Differential diagnosis generation with ranked probability estimates.
|
||||
\item Drug interaction checking with severity classification.
|
||||
\item Clinical note summarization (SOAP format, discharge summaries).
|
||||
\item Laboratory result interpretation with reference ranges.
|
||||
\item Medical imaging description (radiology, pathology, dermatology).
|
||||
\item Patient education materials at Flesch-Kincaid grade 8 reading level.
|
||||
\end{itemize}
|
||||
|
||||
\subsection{Constitutional Medical Alignment}
|
||||
|
||||
Medical AI requires conservative alignment beyond standard RLHF. We apply constitutional AI with medical-specific constraints:
|
||||
|
||||
\begin{itemize}
|
||||
\item \textbf{Uncertainty disclosure}: The model must express calibrated uncertainty for diagnoses and treatments.
|
||||
\item \textbf{Escalation triggers}: Emergency symptoms (chest pain, stroke signs, suicidality) always trigger urgent referral language.
|
||||
\item \textbf{Liability boundaries}: The model disclaims diagnostic authority and frames outputs as clinical decision support.
|
||||
\item \textbf{Source citation}: Clinical claims without citations are penalized during RLHF.
|
||||
\end{itemize}
|
||||
|
||||
\section{Clinical Reasoning Chains}
|
||||
|
||||
\subsection{Chain-of-Thought Clinical Reasoning}
|
||||
|
||||
Zen Medical generates explicit reasoning chains before clinical conclusions. For a differential diagnosis query, the chain follows:
|
||||
|
||||
\begin{enumerate}
|
||||
\item \textbf{Symptom enumeration}: List presenting symptoms with onset, duration, severity, and aggravating/relieving factors.
|
||||
\item \textbf{Hypothesis generation}: Generate candidate diagnoses with prior probability estimates.
|
||||
\item \textbf{Discriminating questions}: Identify clinical features that would distinguish among candidates.
|
||||
\item \textbf{Evidence weighing}: Update hypothesis probabilities based on available findings.
|
||||
\item \textbf{Recommendation}: Provide diagnostic workup and management recommendations with citations.
|
||||
\end{enumerate}
|
||||
|
||||
This structure mirrors the clinical reasoning process taught in medical education, facilitating adoption by clinicians who can validate each step.
|
||||
|
||||
\subsection{Evidence Citation System}
|
||||
|
||||
Citations are generated using a retrieval-augmented generation (RAG) pipeline over a 48-million-article PubMed index combined with a curated guideline database:
|
||||
|
||||
\begin{equation}
|
||||
\text{answer} = \text{ZenMedical}\left(q, \text{TopK}\!\left(\text{Retrieve}(q), K=8\right)\right)
|
||||
\end{equation}
|
||||
|
||||
Citation accuracy (retrieved article actually supports the claim) is evaluated at 94.2\% by clinical expert review.
|
||||
|
||||
\section{Benchmark Results}
|
||||
|
||||
\subsection{MedQA (USMLE Format)}
|
||||
|
||||
\begin{table}[H]
|
||||
\centering
|
||||
\caption{MedQA benchmark results by step}
|
||||
\label{tab:medqa}
|
||||
\begin{tabular}{lcccc}
|
||||
\toprule
|
||||
Benchmark & Questions & Zen Medical & Human Expert Avg & Passing Score \\
|
||||
\midrule
|
||||
USMLE Step 1 & 1,273 & 91.2\% & 82.4\% & 60\% \\
|
||||
USMLE Step 2 CK & 1,048 & 89.8\% & 85.2\% & 60\% \\
|
||||
USMLE Step 3 & 384 & 86.8\% & 83.8\% & 60\% \\
|
||||
\midrule
|
||||
\textbf{MedQA 4-option} & \textbf{11,450} & \textbf{89.3\%} & \textbf{83.8\%} & \textbf{60\%} \\
|
||||
\midrule
|
||||
USMLE Step 1 (our eval) & 2,484 & 88.4\% & 82.8\% & 60\% \\
|
||||
USMLE Step 2 CK & 2,128 & 87.2\% & 85.4\% & 60\% \\
|
||||
\midrule
|
||||
\textbf{USMLE Overall} & \textbf{4,612} & \textbf{87.2\%} & \textbf{84.1\%} & \textbf{60\%} \\
|
||||
\bottomrule
|
||||
\end{tabular}
|
||||
\end{table}
|
||||
|
||||
\subsection{PubMedQA}
|
||||
|
||||
PubMedQA tests biomedical research question answering from abstract context.
|
||||
|
||||
\begin{table}[H]
|
||||
\centering
|
||||
\caption{PubMedQA results}
|
||||
\label{tab:pubmedqa}
|
||||
\begin{tabular}{lccc}
|
||||
\toprule
|
||||
Subset & Questions & Accuracy & F1 \\
|
||||
\midrule
|
||||
Labeled (yes/no/maybe) & 1,000 & 76.4\% & 0.742 \\
|
||||
Unlabeled (few-shot) & 61,249 & 72.8\% & 0.714 \\
|
||||
\bottomrule
|
||||
\end{tabular}
|
||||
\end{table}
|
||||
|
||||
\subsection{Clinical NLP Tasks}
|
||||
|
||||
\begin{table}[H]
|
||||
\centering
|
||||
\caption{Clinical NLP benchmark results}
|
||||
\label{tab:clinical_nlp}
|
||||
\begin{tabular}{llcc}
|
||||
\toprule
|
||||
Task & Benchmark & Metric & Score \\
|
||||
\midrule
|
||||
Clinical NER & i2b2 2012 & F1 & 0.918 \\
|
||||
Medical relation extraction & i2b2 2010 & F1 & 0.872 \\
|
||||
Temporal event extraction & THYME & F1 & 0.841 \\
|
||||
Clinical coreference & i2b2 2011 & F1 & 0.894 \\
|
||||
Diagnosis coding (ICD-10) & MIMIC-III & Micro-F1 & 0.684 \\
|
||||
Medication event extraction & n2c2 2018 & F1 & 0.924 \\
|
||||
Social determinant extraction & n2c2 2022 & F1 & 0.881 \\
|
||||
\bottomrule
|
||||
\end{tabular}
|
||||
\end{table}
|
||||
|
||||
\subsection{Medical Imaging Understanding}
|
||||
|
||||
When integrated with the Zen Vision Architecture:
|
||||
|
||||
\begin{table}[H]
|
||||
\centering
|
||||
\caption{Medical imaging QA and classification}
|
||||
\label{tab:medical_imaging}
|
||||
\begin{tabular}{lcc}
|
||||
\toprule
|
||||
Task & Benchmark & Score \\
|
||||
\midrule
|
||||
Chest X-ray classification & CheXpert (AUC avg) & 0.924 \\
|
||||
Radiology report generation & MIMIC-CXR & BLEU-4 0.184 \\
|
||||
Pathology image QA & PathVQA & 88.2\% \\
|
||||
Dermatology classification & ISIC 2020 & AUC 0.961 \\
|
||||
Ophthalmology (DR grading) & EyePACS & QWK 0.912 \\
|
||||
\bottomrule
|
||||
\end{tabular}
|
||||
\end{table}
|
||||
|
||||
\section{HIPAA-Compliant Deployment}
|
||||
|
||||
\subsection{Architecture Overview}
|
||||
|
||||
Zen Medical is deployed entirely on-premise within the customer's secure network boundary. No patient data or query content leaves the deployment environment.
|
||||
|
||||
\begin{itemize}
|
||||
\item \textbf{Compute}: Customer-managed Kubernetes cluster on bare-metal or private cloud GPU nodes.
|
||||
\item \textbf{Network}: Air-gapped deployment option with no external API calls; fully offline capable.
|
||||
\item \textbf{Storage}: Encrypted at rest (AES-256) and in transit (TLS 1.3 minimum).
|
||||
\item \textbf{Access control}: Role-based access integrated with existing hospital directory services (LDAP/SAML).
|
||||
\item \textbf{Audit logging}: All queries and responses logged with user, timestamp, and PHI-scrubbed content.
|
||||
\end{itemize}
|
||||
|
||||
\subsection{PHI Detection and Redaction}
|
||||
|
||||
Queries and responses pass through a PHI detection layer before logging:
|
||||
\begin{itemize}
|
||||
\item Named entity recognition for person names, dates, locations, phone numbers, MRNs.
|
||||
\item Regex patterns for structured PHI (SSN, NPI, DEA numbers).
|
||||
\item Contextual classifier for ambiguous PHI (e.g., initials in clinical context).
|
||||
\end{itemize}
|
||||
|
||||
PHI recall (fraction of actual PHI detected): 99.94\%. PHI precision: 98.8\%.
|
||||
|
||||
\section{Safety Guardrails}
|
||||
|
||||
\subsection{Uncertainty Quantification}
|
||||
|
||||
Zen Medical outputs calibrated confidence estimates alongside clinical recommendations. Confidence is derived from:
|
||||
|
||||
\begin{equation}
|
||||
\text{conf}(y \mid x) = \mathbb{E}_{k \sim p(\text{experts})} P(y \mid x, \text{expert}_k)
|
||||
\end{equation}
|
||||
|
||||
When confidence falls below threshold $\tau = 0.65$, the system appends a mandatory uncertainty disclosure and suggests specialist consultation.
|
||||
|
||||
\subsection{Emergency Detection}
|
||||
|
||||
A dedicated emergency detection classifier runs on every query with 8ms latency:
|
||||
|
||||
\begin{table}[H]
|
||||
\centering
|
||||
\caption{Emergency detection classifier performance}
|
||||
\label{tab:emergency}
|
||||
\begin{tabular}{lcc}
|
||||
\toprule
|
||||
Emergency Category & Sensitivity & Specificity \\
|
||||
\midrule
|
||||
Chest pain / MI & 99.8\% & 97.4\% \\
|
||||
Stroke symptoms & 99.6\% & 96.8\% \\
|
||||
Suicidal ideation & 99.4\% & 94.2\% \\
|
||||
Severe allergic reaction & 99.2\% & 97.8\% \\
|
||||
Sepsis indicators & 98.8\% & 95.4\% \\
|
||||
\bottomrule
|
||||
\end{tabular}
|
||||
\end{table}
|
||||
|
||||
\section{EHR Integration}
|
||||
|
||||
Zen Medical integrates with major EHR platforms via HL7 FHIR R4 APIs, supporting:
|
||||
\begin{itemize}
|
||||
\item Patient context injection: relevant history, medications, allergies pre-loaded into model context.
|
||||
\item Note generation: SOAP-format clinical notes exported directly to EHR encounter.
|
||||
\item Order suggestions: CPOE-compatible medication and lab order drafts.
|
||||
\item Coding assistance: ICD-10-CM and CPT code suggestions from encounter notes.
|
||||
\end{itemize}
|
||||
|
||||
\section{Conclusion}
|
||||
|
||||
Zen Medical achieves clinician-level performance on standardized medical benchmarks (89.3\% MedQA, 87.2\% USMLE, 0.918 clinical NER F1) while satisfying healthcare's stringent privacy, safety, and integration requirements. Evidence-grounded generation, explicit reasoning chains, and conservative uncertainty disclosure address the trust requirements that historically limited AI adoption in clinical settings. HIPAA-compliant on-premise deployment eliminates data governance barriers for hospital deployment.
|
||||
|
||||
\begin{thebibliography}{99}
|
||||
\bibitem{medpalm} Singhal, K. et al. Large Language Models Encode Clinical Knowledge. \textit{Nature}, 2023.
|
||||
\bibitem{medqa} Jin, D. et al. What Disease Does This Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams. \textit{Applied Sciences}, 2021.
|
||||
\bibitem{pubmedqa} Jin, Q. et al. PubMedQA: A Dataset for Biomedical Research Question Answering. \textit{EMNLP}, 2019.
|
||||
\bibitem{hipaa} U.S. Department of Health and Human Services. HIPAA Security Rule Technical Safeguards. 2013.
|
||||
\bibitem{fhir} HL7 International. Fast Healthcare Interoperability Resources (FHIR) R4. 2019.
|
||||
\end{thebibliography}
|
||||
|
||||
\end{document}
|
||||
Binary file not shown.
+1
-1
@@ -101,7 +101,7 @@
|
||||
}
|
||||
|
||||
\author{
|
||||
\textbf{Hanzo AI Research}$^{1}$ \quad \textbf{Zoo Labs Foundation}$^{2}$ \\[0.6em]
|
||||
Antje Worring, \textbf{Hanzo AI Research}$^{1}$ \quad \textbf{Zoo Labs Foundation}$^{2}$ \\[0.6em]
|
||||
$^{1}$Hanzo AI Inc. (Techstars '17) \quad $^{2}$Zoo Labs Foundation (501(c)(3)) \\[0.3em]
|
||||
\texttt{research@hanzo.ai} \quad \texttt{foundation@zoo.ngo} \\[0.3em]
|
||||
{\small \url{https://hanzo.ai/research/zen-medical}}
|
||||
|
||||
Binary file not shown.
@@ -15,7 +15,7 @@
|
||||
|
||||
\title{\textbf{Zen MoDE: Mixture of Distilled Experts Architecture}\\
|
||||
\large Technical Report v2025.06}
|
||||
\author{Zen LM Research Team\\
|
||||
\author{Antje Worring, Zach Kelling \\ Zen LM Research Team\\
|
||||
\texttt{research@zenlm.org}}
|
||||
\date{June 2025}
|
||||
|
||||
|
||||
Binary file not shown.
@@ -15,7 +15,7 @@
|
||||
|
||||
\title{\textbf{Zen Multilingual: 100+ Language Coverage and Low-Resource NLP}\\
|
||||
\large Technical Report v2025.07}
|
||||
\author{Zen LM Research Team\\
|
||||
\author{Antje Worring, Zach Kelling \\ Zen LM Research Team\\
|
||||
\texttt{research@zenlm.org}}
|
||||
\date{July 2025}
|
||||
|
||||
|
||||
Binary file not shown.
@@ -17,7 +17,7 @@
|
||||
|
||||
\title{\textbf{Zen Multimodal Architecture: Unified Vision-Language-Audio}\\
|
||||
\large Technical Report v2025.03}
|
||||
\author{Zen LM Research Team\\
|
||||
\author{Antje Worring, Zach Kelling \\ Zen LM Research Team\\
|
||||
\texttt{research@zenlm.org}}
|
||||
\date{March 2025}
|
||||
|
||||
@@ -265,7 +265,7 @@ Gated cross-attention (ours) & \textbf{82.4} & \textbf{85.1} \\
|
||||
\bottomrule
|
||||
\end{tabular}
|
||||
\caption{Fusion strategy ablation. Gated cross-attention best preserves text while gaining vision.}
|
||||
\end{table>
|
||||
\end{table}
|
||||
|
||||
\subsection{Effect of Modal Dropout Rate}
|
||||
|
||||
|
||||
Binary file not shown.
+3
-3
@@ -5,7 +5,7 @@
|
||||
\usepackage{graphicx}
|
||||
\usepackage{hyperref}
|
||||
\usepackage{listings}
|
||||
\usepackage{color}
|
||||
\usepackage[dvipsnames]{xcolor}
|
||||
\usepackage{booktabs}
|
||||
\usepackage{float}
|
||||
\usepackage{geometry}
|
||||
@@ -39,7 +39,7 @@
|
||||
}
|
||||
|
||||
\author{
|
||||
Hanzo AI Research Team\thanks{research@hanzo.ai} \and
|
||||
Antje Worring \and Hanzo AI Research Team\thanks{research@hanzo.ai} \and
|
||||
Zoo Labs Foundation\thanks{foundation@zoo.ngo}
|
||||
}
|
||||
|
||||
@@ -233,7 +233,7 @@ Human composers (reference) & 1.8 & -- \\
|
||||
\bottomrule
|
||||
\end{tabular}
|
||||
\caption{Frechet Audio Distance on MusicCaps evaluation set}
|
||||
\end{table>
|
||||
\end{table}
|
||||
|
||||
\subsection{Music Opinion Score (MOS)}
|
||||
|
||||
|
||||
@@ -1,323 +0,0 @@
|
||||
\documentclass[11pt,a4paper]{article}
|
||||
\usepackage[utf8]{inputenc}
|
||||
\usepackage[T1]{fontenc}
|
||||
\usepackage{amsmath,amsfonts,amssymb}
|
||||
\usepackage{graphicx}
|
||||
\usepackage{hyperref}
|
||||
\usepackage{listings}
|
||||
\usepackage{color}
|
||||
\usepackage{booktabs}
|
||||
\usepackage{float}
|
||||
\usepackage{geometry}
|
||||
\geometry{margin=1in}
|
||||
|
||||
% Color definitions
|
||||
\definecolor{zenblue}{RGB}{41,121,255}
|
||||
\definecolor{zengreen}{RGB}{52,199,89}
|
||||
\definecolor{zenorange}{RGB}{255,149,0}
|
||||
\definecolor{codegray}{RGB}{245,245,245}
|
||||
|
||||
% Hyperref setup
|
||||
\hypersetup{
|
||||
colorlinks=true,
|
||||
linkcolor=zenblue,
|
||||
urlcolor=zenblue,
|
||||
citecolor=zenblue
|
||||
}
|
||||
|
||||
% Code listing setup
|
||||
\lstset{
|
||||
backgroundcolor=\color{codegray},
|
||||
basicstyle=\ttfamily\small,
|
||||
breaklines=true,
|
||||
captionpos=b,
|
||||
frame=single,
|
||||
numbers=left,
|
||||
numberstyle=\tiny\color{gray}
|
||||
}
|
||||
|
||||
\title{
|
||||
\vspace{-2cm}
|
||||
\Large \textbf{Zen AI Model Family} \\
|
||||
\vspace{0.5cm}
|
||||
\Huge \textbf{Zen-Nano} \\
|
||||
\vspace{0.3cm}
|
||||
\large Mobile/IoT Intelligence \\
|
||||
\vspace{0.5cm}
|
||||
\normalsize Technical Whitepaper v1.0
|
||||
}
|
||||
|
||||
\author{
|
||||
Zach Kelling\thanks{zach@lux.network} \\
|
||||
\texttt{research@hanzo.ai} \\
|
||||
\\
|
||||
Zoo Labs Foundation \\
|
||||
\texttt{foundation@zoolabs.org}
|
||||
}
|
||||
|
||||
\date{September 2025}
|
||||
|
||||
\begin{document}
|
||||
|
||||
\maketitle
|
||||
|
||||
\begin{abstract}
|
||||
We present \textbf{Zen-Nano}, a 0.6B parameter model optimized for mobile/iot intelligence.
|
||||
Built upon zen-0.5B, this model achieves state-of-the-art performance while maintaining exceptional efficiency
|
||||
with only 0.6B active parameters. Supporting 64K thinking tokens for advanced reasoning, the model represents a significant advancement in democratizing AI through sustainable and efficient architectures.
|
||||
\end{abstract}
|
||||
|
||||
\tableofcontents
|
||||
\newpage
|
||||
|
||||
\section{Introduction}
|
||||
|
||||
The rapid advancement of artificial intelligence has created an unprecedented demand for models that balance capability with efficiency.
|
||||
\textbf{Zen-Nano} addresses this challenge by delivering enterprise-grade performance while maintaining a minimal computational footprint.
|
||||
|
||||
\subsection{Key Innovations}
|
||||
\begin{itemize}
|
||||
\item \textbf{Efficient Architecture}: 0.6B active parameters from 0.6B total
|
||||
\item \textbf{Specialized Training}: Optimized for mobile/iot intelligence
|
||||
\item \textbf{Extended Context}: 32K context window
|
||||
\item \textbf{Thinking Mode}: 64K thinking tokens
|
||||
|
||||
|
||||
\end{itemize}
|
||||
|
||||
\section{Architecture}
|
||||
|
||||
\subsection{Model Design}
|
||||
|
||||
Zen-Nano is based on the zen-0.5B architecture with several key modifications:
|
||||
|
||||
\begin{table}[H]
|
||||
\centering
|
||||
\begin{tabular}{ll}
|
||||
\toprule
|
||||
\textbf{Component} & \textbf{Specification} \\
|
||||
\midrule
|
||||
Total Parameters & 0.6B \\
|
||||
Active Parameters & 0.6B \\
|
||||
Base Model & zen-0.5B \\
|
||||
Context Length & 32K \\
|
||||
Thinking Tokens & 64K \\
|
||||
|
||||
|
||||
Architecture Type & Transformer \\
|
||||
\bottomrule
|
||||
\end{tabular}
|
||||
\caption{Zen-Nano Architecture Specifications}
|
||||
\end{table}
|
||||
|
||||
\subsection{Technical Innovations}
|
||||
|
||||
\subsubsection{Mixture of Experts (MoE)}
|
||||
The model uses a dense architecture with all parameters active during inference, optimized for maximum performance per parameter.
|
||||
|
||||
\subsubsection{Attention Mechanism}
|
||||
Extended attention mechanisms support up to 32K context length with efficient KV-cache management.
|
||||
|
||||
\subsubsection{Thinking Mode}
|
||||
Advanced reasoning through extended thinking tokens (up to 64K), enabling:
|
||||
\begin{itemize}
|
||||
\item Step-by-step problem decomposition
|
||||
\item Self-correction and verification
|
||||
\item Complex multi-step reasoning
|
||||
\item Internal deliberation before response
|
||||
\end{itemize}
|
||||
|
||||
\section{Performance Benchmarks}
|
||||
|
||||
\subsection{Evaluation Results}
|
||||
|
||||
|
||||
\begin{table}[H]
|
||||
\centering
|
||||
\begin{tabular}{lc}
|
||||
\toprule
|
||||
\textbf{Benchmark} & \textbf{Score} \\
|
||||
\midrule
|
||||
MMLU & 51.7\% \\
|
||||
HumanEval & 22.6\% \\
|
||||
GSM8K & 62.0\% \\
|
||||
HellaSwag & 59.5\% \\
|
||||
\bottomrule
|
||||
\end{tabular}
|
||||
\caption{Language Understanding Benchmarks}
|
||||
\end{table}
|
||||
|
||||
\subsection{Efficiency Metrics}
|
||||
|
||||
\begin{table}[H]
|
||||
\centering
|
||||
\begin{tabular}{ll}
|
||||
\toprule
|
||||
\textbf{Metric} & \textbf{Value} \\
|
||||
\midrule
|
||||
Inference Speed & 450 tokens/sec \\
|
||||
Memory Usage (INT4) & 2 GB \\
|
||||
Energy Efficiency & 98\% reduction \\
|
||||
Latency (First Token) & 15 ms \\
|
||||
\bottomrule
|
||||
\end{tabular}
|
||||
\caption{Efficiency Metrics}
|
||||
\end{table}
|
||||
|
||||
\section{Training Methodology}
|
||||
|
||||
\subsection{Dataset}
|
||||
The model was trained on a carefully curated dataset comprising:
|
||||
\begin{itemize}
|
||||
\item High-quality filtered web data (0.5TB)
|
||||
\item Domain-specific corpora for mobile/iot intelligence
|
||||
\item Synthetic data generation for edge cases
|
||||
\item Human feedback through RLHF
|
||||
\end{itemize}
|
||||
|
||||
\subsection{Training Process}
|
||||
\begin{enumerate}
|
||||
\item \textbf{Pretraining}: 2 trillion tokens over 14 days on 8x A100
|
||||
\item \textbf{Supervised Fine-tuning}: Task-specific optimization
|
||||
\item \textbf{RLHF}: Alignment with human preferences
|
||||
\item \textbf{Constitutional AI}: Safety and helpfulness optimization
|
||||
\end{enumerate}
|
||||
|
||||
\section{Use Cases and Applications}
|
||||
|
||||
\subsection{Primary Applications}
|
||||
\item Conversational AI and chatbots
|
||||
\item Content generation and summarization
|
||||
\item Code completion and review
|
||||
\item Educational assistance
|
||||
\item Research and analysis
|
||||
|
||||
\subsection{Integration Examples}
|
||||
|
||||
\begin{lstlisting}[language=Python, caption=Basic Usage Example]
|
||||
from transformers import AutoModelForCausalLM, AutoTokenizer
|
||||
|
||||
# Load model and tokenizer
|
||||
model = AutoModelForCausalLM.from_pretrained("zenlm/zen-nano-0.6b-instruct")
|
||||
tokenizer = AutoTokenizer.from_pretrained("zenlm/zen-nano-0.6b-instruct")
|
||||
|
||||
# Generate response
|
||||
inputs = tokenizer("Explain quantum computing", return_tensors="pt")
|
||||
outputs = model.generate(**inputs, max_length=100)
|
||||
response = tokenizer.decode(outputs[0])
|
||||
\end{lstlisting}
|
||||
|
||||
\section{Environmental Impact}
|
||||
|
||||
\subsection{Sustainability Metrics}
|
||||
\begin{itemize}
|
||||
\item \textbf{Carbon Footprint}: 0.02 kg CO₂e per million inferences
|
||||
\item \textbf{Energy Usage}: 0.5 kWh per day (1000 users)
|
||||
\item \textbf{Efficiency Gain}: 98\% reduction vs comparable models
|
||||
\end{itemize}
|
||||
|
||||
\subsection{Green AI Commitment}
|
||||
Zen AI models are designed with sustainability as a core principle, achieving industry-leading efficiency
|
||||
through architectural innovations and optimization techniques.
|
||||
|
||||
\section{Safety and Alignment}
|
||||
|
||||
\subsection{Safety Measures}
|
||||
\begin{itemize}
|
||||
\item Constitutional AI training for harmlessness
|
||||
\item Comprehensive red-teaming and adversarial testing
|
||||
\item Built-in safety filters and guardrails
|
||||
\item Regular safety audits and updates
|
||||
\end{itemize}
|
||||
|
||||
\subsection{Ethical Considerations}
|
||||
The model has been developed with careful attention to:
|
||||
\begin{itemize}
|
||||
\item Bias mitigation through diverse training data
|
||||
\item Transparency in capabilities and limitations
|
||||
\item Privacy-preserving deployment options
|
||||
\item Responsible AI principles alignment
|
||||
\end{itemize}
|
||||
|
||||
\section{Deployment Options}
|
||||
|
||||
\subsection{Available Formats}
|
||||
\begin{itemize}
|
||||
\item \textbf{SafeTensors}: Original precision weights
|
||||
\item \textbf{GGUF}: Quantized formats (Q4\_K\_M, Q5\_K\_M, Q8\_0)
|
||||
\item \textbf{MLX}: Apple Silicon optimization (4-bit, 8-bit)
|
||||
\item \textbf{ONNX}: Cross-platform deployment (coming soon)
|
||||
\end{itemize}
|
||||
|
||||
\subsection{Hardware Requirements}
|
||||
\begin{table}[H]
|
||||
\centering
|
||||
\begin{tabular}{lll}
|
||||
\toprule
|
||||
\textbf{Precision} & \textbf{Memory} & \textbf{Recommended Hardware} \\
|
||||
\midrule
|
||||
FP16 & 1.2 GB & RTX 3060 \\
|
||||
INT8 & 0.6 GB & GTX 1660 \\
|
||||
INT4 & 2 GB & Raspberry Pi 5 \\
|
||||
\bottomrule
|
||||
\end{tabular}
|
||||
\caption{Hardware Requirements by Precision}
|
||||
\end{table}
|
||||
|
||||
\section{Future Work}
|
||||
|
||||
\subsection{Planned Improvements}
|
||||
\begin{itemize}
|
||||
\item Extended context windows (up to 1M tokens)
|
||||
\item Enhanced multimodal capabilities
|
||||
\item Improved efficiency through further optimization
|
||||
\item Expanded language support
|
||||
\end{itemize}
|
||||
|
||||
\subsection{Research Directions}
|
||||
\begin{itemize}
|
||||
\item Advanced reasoning mechanisms
|
||||
\item Self-supervised learning improvements
|
||||
\item Zero-shot generalization enhancement
|
||||
\item Continual learning capabilities
|
||||
\end{itemize}
|
||||
|
||||
\section{Conclusion}
|
||||
|
||||
\textbf{Zen-Nano} represents a significant advancement in AI democratization,
|
||||
delivering exceptional performance for mobile/iot intelligence while maintaining
|
||||
unprecedented efficiency. Through innovative architecture design and careful optimization,
|
||||
the model achieves a balance between capability and sustainability that sets a new standard
|
||||
for responsible AI development.
|
||||
|
||||
\section*{Acknowledgments}
|
||||
|
||||
We thank the open-source community, our research partners, and the teams at Hanzo AI and
|
||||
Zoo Labs Foundation for their contributions to this work.
|
||||
|
||||
\bibliographystyle{plain}
|
||||
\bibliography{references}
|
||||
|
||||
\appendix
|
||||
|
||||
\section{Model Card}
|
||||
|
||||
\begin{table}[H]
|
||||
\centering
|
||||
\begin{tabular}{ll}
|
||||
\toprule
|
||||
\textbf{Field} & \textbf{Value} \\
|
||||
\midrule
|
||||
Model Name & Zen-Nano \\
|
||||
Version & 1.0.0 \\
|
||||
Release Date & September 2025 \\
|
||||
License & Apache 2.0 \\
|
||||
Repository & \href{https://huggingface.co/zenlm/zen-nano-0.6b-instruct}{huggingface.co/zenlm/zen-nano-0.6b-instruct} \\
|
||||
Documentation & \href{https://github.com/zenlm/zen}{github.com/zenlm/zen} \\
|
||||
Contact & research@hanzo.ai \\
|
||||
\bottomrule
|
||||
\end{tabular}
|
||||
\caption{Model Card Information}
|
||||
\end{table}
|
||||
|
||||
\end{document}
|
||||
@@ -1,323 +0,0 @@
|
||||
\documentclass[11pt,a4paper]{article}
|
||||
\usepackage[utf8]{inputenc}
|
||||
\usepackage[T1]{fontenc}
|
||||
\usepackage{amsmath,amsfonts,amssymb}
|
||||
\usepackage{graphicx}
|
||||
\usepackage{hyperref}
|
||||
\usepackage{listings}
|
||||
\usepackage{color}
|
||||
\usepackage{booktabs}
|
||||
\usepackage{float}
|
||||
\usepackage{geometry}
|
||||
\geometry{margin=1in}
|
||||
|
||||
% Color definitions
|
||||
\definecolor{zenblue}{RGB}{41,121,255}
|
||||
\definecolor{zengreen}{RGB}{52,199,89}
|
||||
\definecolor{zenorange}{RGB}{255,149,0}
|
||||
\definecolor{codegray}{RGB}{245,245,245}
|
||||
|
||||
% Hyperref setup
|
||||
\hypersetup{
|
||||
colorlinks=true,
|
||||
linkcolor=zenblue,
|
||||
urlcolor=zenblue,
|
||||
citecolor=zenblue
|
||||
}
|
||||
|
||||
% Code listing setup
|
||||
\lstset{
|
||||
backgroundcolor=\color{codegray},
|
||||
basicstyle=\ttfamily\small,
|
||||
breaklines=true,
|
||||
captionpos=b,
|
||||
frame=single,
|
||||
numbers=left,
|
||||
numberstyle=\tiny\color{gray}
|
||||
}
|
||||
|
||||
\title{
|
||||
\vspace{-2cm}
|
||||
\Large \textbf{Zen AI Model Family} \\
|
||||
\vspace{0.5cm}
|
||||
\Huge \textbf{Zen-Next} \\
|
||||
\vspace{0.3cm}
|
||||
\large Flagship Model \\
|
||||
\vspace{0.5cm}
|
||||
\normalsize Technical Whitepaper v1.0
|
||||
}
|
||||
|
||||
\author{
|
||||
Zach Kelling\thanks{zach@lux.network} \\
|
||||
\texttt{research@hanzo.ai} \\
|
||||
\\
|
||||
Zoo Labs Foundation \\
|
||||
\texttt{foundation@zoolabs.org}
|
||||
}
|
||||
|
||||
\date{September 2025}
|
||||
|
||||
\begin{document}
|
||||
|
||||
\maketitle
|
||||
|
||||
\begin{abstract}
|
||||
We present \textbf{Zen-Next}, a 80B parameter model optimized for flagship model.
|
||||
Built upon zen-72B, this model achieves state-of-the-art performance while maintaining exceptional efficiency
|
||||
with only 80B active parameters. Supporting 1M thinking tokens for advanced reasoning, the model represents a significant advancement in democratizing AI through sustainable and efficient architectures.
|
||||
\end{abstract}
|
||||
|
||||
\tableofcontents
|
||||
\newpage
|
||||
|
||||
\section{Introduction}
|
||||
|
||||
The rapid advancement of artificial intelligence has created an unprecedented demand for models that balance capability with efficiency.
|
||||
\textbf{Zen-Next} addresses this challenge by delivering enterprise-grade performance while maintaining a minimal computational footprint.
|
||||
|
||||
\subsection{Key Innovations}
|
||||
\begin{itemize}
|
||||
\item \textbf{Efficient Architecture}: 80B active parameters from 80B total
|
||||
\item \textbf{Specialized Training}: Optimized for flagship model
|
||||
\item \textbf{Extended Context}: 128K context window
|
||||
\item \textbf{Thinking Mode}: 1M thinking tokens
|
||||
|
||||
|
||||
\end{itemize}
|
||||
|
||||
\section{Architecture}
|
||||
|
||||
\subsection{Model Design}
|
||||
|
||||
Zen-Next is based on the zen-72B architecture with several key modifications:
|
||||
|
||||
\begin{table}[H]
|
||||
\centering
|
||||
\begin{tabular}{ll}
|
||||
\toprule
|
||||
\textbf{Component} & \textbf{Specification} \\
|
||||
\midrule
|
||||
Total Parameters & 80B \\
|
||||
Active Parameters & 80B \\
|
||||
Base Model & zen-72B \\
|
||||
Context Length & 128K \\
|
||||
Thinking Tokens & 1M \\
|
||||
|
||||
|
||||
Architecture Type & Transformer \\
|
||||
\bottomrule
|
||||
\end{tabular}
|
||||
\caption{Zen-Next Architecture Specifications}
|
||||
\end{table}
|
||||
|
||||
\subsection{Technical Innovations}
|
||||
|
||||
\subsubsection{Mixture of Experts (MoE)}
|
||||
The model uses a dense architecture with all parameters active during inference, optimized for maximum performance per parameter.
|
||||
|
||||
\subsubsection{Attention Mechanism}
|
||||
Extended attention mechanisms support up to 128K context length with efficient KV-cache management.
|
||||
|
||||
\subsubsection{Thinking Mode}
|
||||
Advanced reasoning through extended thinking tokens (up to 1M), enabling:
|
||||
\begin{itemize}
|
||||
\item Step-by-step problem decomposition
|
||||
\item Self-correction and verification
|
||||
\item Complex multi-step reasoning
|
||||
\item Internal deliberation before response
|
||||
\end{itemize}
|
||||
|
||||
\section{Performance Benchmarks}
|
||||
|
||||
\subsection{Evaluation Results}
|
||||
|
||||
|
||||
\begin{table}[H]
|
||||
\centering
|
||||
\begin{tabular}{lc}
|
||||
\toprule
|
||||
\textbf{Benchmark} & \textbf{Score} \\
|
||||
\midrule
|
||||
MMLU & 75.6\% \\
|
||||
HumanEval & 61.7\% \\
|
||||
GSM8K & 90.7\% \\
|
||||
HellaSwag & 86.9\% \\
|
||||
\bottomrule
|
||||
\end{tabular}
|
||||
\caption{Language Understanding Benchmarks}
|
||||
\end{table}
|
||||
|
||||
\subsection{Efficiency Metrics}
|
||||
|
||||
\begin{table}[H]
|
||||
\centering
|
||||
\begin{tabular}{ll}
|
||||
\toprule
|
||||
\textbf{Metric} & \textbf{Value} \\
|
||||
\midrule
|
||||
Inference Speed & 45 tokens/sec \\
|
||||
Memory Usage (INT4) & 40 GB \\
|
||||
Energy Efficiency & 80\% reduction \\
|
||||
Latency (First Token) & 200 ms \\
|
||||
\bottomrule
|
||||
\end{tabular}
|
||||
\caption{Efficiency Metrics}
|
||||
\end{table}
|
||||
|
||||
\section{Training Methodology}
|
||||
|
||||
\subsection{Dataset}
|
||||
The model was trained on a carefully curated dataset comprising:
|
||||
\begin{itemize}
|
||||
\item High-quality filtered web data (35TB)
|
||||
\item Domain-specific corpora for flagship model
|
||||
\item Synthetic data generation for edge cases
|
||||
\item Human feedback through RLHF
|
||||
\end{itemize}
|
||||
|
||||
\subsection{Training Process}
|
||||
\begin{enumerate}
|
||||
\item \textbf{Pretraining}: 5 trillion tokens over 45 days on 64x A100
|
||||
\item \textbf{Supervised Fine-tuning}: Task-specific optimization
|
||||
\item \textbf{RLHF}: Alignment with human preferences
|
||||
\item \textbf{Constitutional AI}: Safety and helpfulness optimization
|
||||
\end{enumerate}
|
||||
|
||||
\section{Use Cases and Applications}
|
||||
|
||||
\subsection{Primary Applications}
|
||||
\item Conversational AI and chatbots
|
||||
\item Content generation and summarization
|
||||
\item Code completion and review
|
||||
\item Educational assistance
|
||||
\item Research and analysis
|
||||
|
||||
\subsection{Integration Examples}
|
||||
|
||||
\begin{lstlisting}[language=Python, caption=Basic Usage Example]
|
||||
from transformers import AutoModelForCausalLM, AutoTokenizer
|
||||
|
||||
# Load model and tokenizer
|
||||
model = AutoModelForCausalLM.from_pretrained("zenlm/zen-next-80b-instruct")
|
||||
tokenizer = AutoTokenizer.from_pretrained("zenlm/zen-next-80b-instruct")
|
||||
|
||||
# Generate response
|
||||
inputs = tokenizer("Explain quantum computing", return_tensors="pt")
|
||||
outputs = model.generate(**inputs, max_length=100)
|
||||
response = tokenizer.decode(outputs[0])
|
||||
\end{lstlisting}
|
||||
|
||||
\section{Environmental Impact}
|
||||
|
||||
\subsection{Sustainability Metrics}
|
||||
\begin{itemize}
|
||||
\item \textbf{Carbon Footprint}: 0.45 kg CO₂e per million inferences
|
||||
\item \textbf{Energy Usage}: 10.2 kWh per day (1000 users)
|
||||
\item \textbf{Efficiency Gain}: 80\% reduction vs comparable models
|
||||
\end{itemize}
|
||||
|
||||
\subsection{Green AI Commitment}
|
||||
Zen AI models are designed with sustainability as a core principle, achieving industry-leading efficiency
|
||||
through architectural innovations and optimization techniques.
|
||||
|
||||
\section{Safety and Alignment}
|
||||
|
||||
\subsection{Safety Measures}
|
||||
\begin{itemize}
|
||||
\item Constitutional AI training for harmlessness
|
||||
\item Comprehensive red-teaming and adversarial testing
|
||||
\item Built-in safety filters and guardrails
|
||||
\item Regular safety audits and updates
|
||||
\end{itemize}
|
||||
|
||||
\subsection{Ethical Considerations}
|
||||
The model has been developed with careful attention to:
|
||||
\begin{itemize}
|
||||
\item Bias mitigation through diverse training data
|
||||
\item Transparency in capabilities and limitations
|
||||
\item Privacy-preserving deployment options
|
||||
\item Responsible AI principles alignment
|
||||
\end{itemize}
|
||||
|
||||
\section{Deployment Options}
|
||||
|
||||
\subsection{Available Formats}
|
||||
\begin{itemize}
|
||||
\item \textbf{SafeTensors}: Original precision weights
|
||||
\item \textbf{GGUF}: Quantized formats (Q4\_K\_M, Q5\_K\_M, Q8\_0)
|
||||
\item \textbf{MLX}: Apple Silicon optimization (4-bit, 8-bit)
|
||||
\item \textbf{ONNX}: Cross-platform deployment (coming soon)
|
||||
\end{itemize}
|
||||
|
||||
\subsection{Hardware Requirements}
|
||||
\begin{table}[H]
|
||||
\centering
|
||||
\begin{tabular}{lll}
|
||||
\toprule
|
||||
\textbf{Precision} & \textbf{Memory} & \textbf{Recommended Hardware} \\
|
||||
\midrule
|
||||
FP16 & 160 GB & 2x A100 80GB \\
|
||||
INT8 & 80 GB & A100 80GB \\
|
||||
INT4 & 40 GB & RTX 4090 \\
|
||||
\bottomrule
|
||||
\end{tabular}
|
||||
\caption{Hardware Requirements by Precision}
|
||||
\end{table}
|
||||
|
||||
\section{Future Work}
|
||||
|
||||
\subsection{Planned Improvements}
|
||||
\begin{itemize}
|
||||
\item Extended context windows (up to 1M tokens)
|
||||
\item Enhanced multimodal capabilities
|
||||
\item Improved efficiency through further optimization
|
||||
\item Expanded language support
|
||||
\end{itemize}
|
||||
|
||||
\subsection{Research Directions}
|
||||
\begin{itemize}
|
||||
\item Advanced reasoning mechanisms
|
||||
\item Self-supervised learning improvements
|
||||
\item Zero-shot generalization enhancement
|
||||
\item Continual learning capabilities
|
||||
\end{itemize}
|
||||
|
||||
\section{Conclusion}
|
||||
|
||||
\textbf{Zen-Next} represents a significant advancement in AI democratization,
|
||||
delivering exceptional performance for flagship model while maintaining
|
||||
unprecedented efficiency. Through innovative architecture design and careful optimization,
|
||||
the model achieves a balance between capability and sustainability that sets a new standard
|
||||
for responsible AI development.
|
||||
|
||||
\section*{Acknowledgments}
|
||||
|
||||
We thank the open-source community, our research partners, and the teams at Hanzo AI and
|
||||
Zoo Labs Foundation for their contributions to this work.
|
||||
|
||||
\bibliographystyle{plain}
|
||||
\bibliography{references}
|
||||
|
||||
\appendix
|
||||
|
||||
\section{Model Card}
|
||||
|
||||
\begin{table}[H]
|
||||
\centering
|
||||
\begin{tabular}{ll}
|
||||
\toprule
|
||||
\textbf{Field} & \textbf{Value} \\
|
||||
\midrule
|
||||
Model Name & Zen-Next \\
|
||||
Version & 1.0.0 \\
|
||||
Release Date & September 2025 \\
|
||||
License & Apache 2.0 \\
|
||||
Repository & \href{https://huggingface.co/zenlm/zen-next-80b-instruct}{huggingface.co/zenlm/zen-next-80b-instruct} \\
|
||||
Documentation & \href{https://github.com/zenlm/zen}{github.com/zenlm/zen} \\
|
||||
Contact & research@hanzo.ai \\
|
||||
\bottomrule
|
||||
\end{tabular}
|
||||
\caption{Model Card Information}
|
||||
\end{table}
|
||||
|
||||
\end{document}
|
||||
@@ -1,323 +0,0 @@
|
||||
\documentclass[11pt,a4paper]{article}
|
||||
\usepackage[utf8]{inputenc}
|
||||
\usepackage[T1]{fontenc}
|
||||
\usepackage{amsmath,amsfonts,amssymb}
|
||||
\usepackage{graphicx}
|
||||
\usepackage{hyperref}
|
||||
\usepackage{listings}
|
||||
\usepackage{color}
|
||||
\usepackage{booktabs}
|
||||
\usepackage{float}
|
||||
\usepackage{geometry}
|
||||
\geometry{margin=1in}
|
||||
|
||||
% Color definitions
|
||||
\definecolor{zenblue}{RGB}{41,121,255}
|
||||
\definecolor{zengreen}{RGB}{52,199,89}
|
||||
\definecolor{zenorange}{RGB}{255,149,0}
|
||||
\definecolor{codegray}{RGB}{245,245,245}
|
||||
|
||||
% Hyperref setup
|
||||
\hypersetup{
|
||||
colorlinks=true,
|
||||
linkcolor=zenblue,
|
||||
urlcolor=zenblue,
|
||||
citecolor=zenblue
|
||||
}
|
||||
|
||||
% Code listing setup
|
||||
\lstset{
|
||||
backgroundcolor=\color{codegray},
|
||||
basicstyle=\ttfamily\small,
|
||||
breaklines=true,
|
||||
captionpos=b,
|
||||
frame=single,
|
||||
numbers=left,
|
||||
numberstyle=\tiny\color{gray}
|
||||
}
|
||||
|
||||
\title{
|
||||
\vspace{-2cm}
|
||||
\Large \textbf{Zen AI Model Family} \\
|
||||
\vspace{0.5cm}
|
||||
\Huge \textbf{Zen-Omni} \\
|
||||
\vspace{0.3cm}
|
||||
\large Multimodal Text \\
|
||||
\vspace{0.5cm}
|
||||
\normalsize Technical Whitepaper v1.0
|
||||
}
|
||||
|
||||
\author{
|
||||
Zach Kelling\thanks{zach@lux.network} \\
|
||||
\texttt{research@hanzo.ai} \\
|
||||
\\
|
||||
Zoo Labs Foundation \\
|
||||
\texttt{foundation@zoolabs.org}
|
||||
}
|
||||
|
||||
\date{September 2025}
|
||||
|
||||
\begin{document}
|
||||
|
||||
\maketitle
|
||||
|
||||
\begin{abstract}
|
||||
We present \textbf{Zen-Omni}, a 30B parameter model optimized for multimodal text.
|
||||
Built upon zen-32B, this model achieves state-of-the-art performance while maintaining exceptional efficiency
|
||||
with only 30B active parameters. Supporting 256K thinking tokens for advanced reasoning, the model represents a significant advancement in democratizing AI through sustainable and efficient architectures.
|
||||
\end{abstract}
|
||||
|
||||
\tableofcontents
|
||||
\newpage
|
||||
|
||||
\section{Introduction}
|
||||
|
||||
The rapid advancement of artificial intelligence has created an unprecedented demand for models that balance capability with efficiency.
|
||||
\textbf{Zen-Omni} addresses this challenge by delivering enterprise-grade performance while maintaining a minimal computational footprint.
|
||||
|
||||
\subsection{Key Innovations}
|
||||
\begin{itemize}
|
||||
\item \textbf{Efficient Architecture}: 30B active parameters from 30B total
|
||||
\item \textbf{Specialized Training}: Optimized for multimodal text
|
||||
\item \textbf{Extended Context}: 128K context window
|
||||
\item \textbf{Thinking Mode}: 256K thinking tokens
|
||||
|
||||
|
||||
\end{itemize}
|
||||
|
||||
\section{Architecture}
|
||||
|
||||
\subsection{Model Design}
|
||||
|
||||
Zen-Omni is based on the zen-32B architecture with several key modifications:
|
||||
|
||||
\begin{table}[H]
|
||||
\centering
|
||||
\begin{tabular}{ll}
|
||||
\toprule
|
||||
\textbf{Component} & \textbf{Specification} \\
|
||||
\midrule
|
||||
Total Parameters & 30B \\
|
||||
Active Parameters & 30B \\
|
||||
Base Model & zen-32B \\
|
||||
Context Length & 128K \\
|
||||
Thinking Tokens & 256K \\
|
||||
|
||||
|
||||
Architecture Type & Transformer \\
|
||||
\bottomrule
|
||||
\end{tabular}
|
||||
\caption{Zen-Omni Architecture Specifications}
|
||||
\end{table}
|
||||
|
||||
\subsection{Technical Innovations}
|
||||
|
||||
\subsubsection{Mixture of Experts (MoE)}
|
||||
The model uses a dense architecture with all parameters active during inference, optimized for maximum performance per parameter.
|
||||
|
||||
\subsubsection{Attention Mechanism}
|
||||
Extended attention mechanisms support up to 128K context length with efficient KV-cache management.
|
||||
|
||||
\subsubsection{Thinking Mode}
|
||||
Advanced reasoning through extended thinking tokens (up to 256K), enabling:
|
||||
\begin{itemize}
|
||||
\item Step-by-step problem decomposition
|
||||
\item Self-correction and verification
|
||||
\item Complex multi-step reasoning
|
||||
\item Internal deliberation before response
|
||||
\end{itemize}
|
||||
|
||||
\section{Performance Benchmarks}
|
||||
|
||||
\subsection{Evaluation Results}
|
||||
|
||||
|
||||
\begin{table}[H]
|
||||
\centering
|
||||
\begin{tabular}{lc}
|
||||
\toprule
|
||||
\textbf{Benchmark} & \textbf{Score} \\
|
||||
\midrule
|
||||
MMLU & 68.4\% \\
|
||||
HumanEval & 48.3\% \\
|
||||
GSM8K & 82.1\% \\
|
||||
HellaSwag & 78.7\% \\
|
||||
\bottomrule
|
||||
\end{tabular}
|
||||
\caption{Language Understanding Benchmarks}
|
||||
\end{table}
|
||||
|
||||
\subsection{Efficiency Metrics}
|
||||
|
||||
\begin{table}[H]
|
||||
\centering
|
||||
\begin{tabular}{ll}
|
||||
\toprule
|
||||
\textbf{Metric} & \textbf{Value} \\
|
||||
\midrule
|
||||
Inference Speed & 85 tokens/sec \\
|
||||
Memory Usage (INT4) & 15 GB \\
|
||||
Energy Efficiency & 85\% reduction \\
|
||||
Latency (First Token) & 120 ms \\
|
||||
\bottomrule
|
||||
\end{tabular}
|
||||
\caption{Efficiency Metrics}
|
||||
\end{table}
|
||||
|
||||
\section{Training Methodology}
|
||||
|
||||
\subsection{Dataset}
|
||||
The model was trained on a carefully curated dataset comprising:
|
||||
\begin{itemize}
|
||||
\item High-quality filtered web data (15TB)
|
||||
\item Domain-specific corpora for multimodal text
|
||||
\item Synthetic data generation for edge cases
|
||||
\item Human feedback through RLHF
|
||||
\end{itemize}
|
||||
|
||||
\subsection{Training Process}
|
||||
\begin{enumerate}
|
||||
\item \textbf{Pretraining}: 5 trillion tokens over 45 days on 64x A100
|
||||
\item \textbf{Supervised Fine-tuning}: Task-specific optimization
|
||||
\item \textbf{RLHF}: Alignment with human preferences
|
||||
\item \textbf{Constitutional AI}: Safety and helpfulness optimization
|
||||
\end{enumerate}
|
||||
|
||||
\section{Use Cases and Applications}
|
||||
|
||||
\subsection{Primary Applications}
|
||||
\item Conversational AI and chatbots
|
||||
\item Content generation and summarization
|
||||
\item Code completion and review
|
||||
\item Educational assistance
|
||||
\item Research and analysis
|
||||
|
||||
\subsection{Integration Examples}
|
||||
|
||||
\begin{lstlisting}[language=Python, caption=Basic Usage Example]
|
||||
from transformers import AutoModelForCausalLM, AutoTokenizer
|
||||
|
||||
# Load model and tokenizer
|
||||
model = AutoModelForCausalLM.from_pretrained("zenlm/zen-omni-30b-instruct")
|
||||
tokenizer = AutoTokenizer.from_pretrained("zenlm/zen-omni-30b-instruct")
|
||||
|
||||
# Generate response
|
||||
inputs = tokenizer("Explain quantum computing", return_tensors="pt")
|
||||
outputs = model.generate(**inputs, max_length=100)
|
||||
response = tokenizer.decode(outputs[0])
|
||||
\end{lstlisting}
|
||||
|
||||
\section{Environmental Impact}
|
||||
|
||||
\subsection{Sustainability Metrics}
|
||||
\begin{itemize}
|
||||
\item \textbf{Carbon Footprint}: 0.25 kg CO₂e per million inferences
|
||||
\item \textbf{Energy Usage}: 5.5 kWh per day (1000 users)
|
||||
\item \textbf{Efficiency Gain}: 85\% reduction vs comparable models
|
||||
\end{itemize}
|
||||
|
||||
\subsection{Green AI Commitment}
|
||||
Zen AI models are designed with sustainability as a core principle, achieving industry-leading efficiency
|
||||
through architectural innovations and optimization techniques.
|
||||
|
||||
\section{Safety and Alignment}
|
||||
|
||||
\subsection{Safety Measures}
|
||||
\begin{itemize}
|
||||
\item Constitutional AI training for harmlessness
|
||||
\item Comprehensive red-teaming and adversarial testing
|
||||
\item Built-in safety filters and guardrails
|
||||
\item Regular safety audits and updates
|
||||
\end{itemize}
|
||||
|
||||
\subsection{Ethical Considerations}
|
||||
The model has been developed with careful attention to:
|
||||
\begin{itemize}
|
||||
\item Bias mitigation through diverse training data
|
||||
\item Transparency in capabilities and limitations
|
||||
\item Privacy-preserving deployment options
|
||||
\item Responsible AI principles alignment
|
||||
\end{itemize}
|
||||
|
||||
\section{Deployment Options}
|
||||
|
||||
\subsection{Available Formats}
|
||||
\begin{itemize}
|
||||
\item \textbf{SafeTensors}: Original precision weights
|
||||
\item \textbf{GGUF}: Quantized formats (Q4\_K\_M, Q5\_K\_M, Q8\_0)
|
||||
\item \textbf{MLX}: Apple Silicon optimization (4-bit, 8-bit)
|
||||
\item \textbf{ONNX}: Cross-platform deployment (coming soon)
|
||||
\end{itemize}
|
||||
|
||||
\subsection{Hardware Requirements}
|
||||
\begin{table}[H]
|
||||
\centering
|
||||
\begin{tabular}{lll}
|
||||
\toprule
|
||||
\textbf{Precision} & \textbf{Memory} & \textbf{Recommended Hardware} \\
|
||||
\midrule
|
||||
FP16 & 60 GB & A100 40GB \\
|
||||
INT8 & 30 GB & RTX 4090 \\
|
||||
INT4 & 15 GB & M2 Max MacBook Pro \\
|
||||
\bottomrule
|
||||
\end{tabular}
|
||||
\caption{Hardware Requirements by Precision}
|
||||
\end{table}
|
||||
|
||||
\section{Future Work}
|
||||
|
||||
\subsection{Planned Improvements}
|
||||
\begin{itemize}
|
||||
\item Extended context windows (up to 1M tokens)
|
||||
\item Enhanced multimodal capabilities
|
||||
\item Improved efficiency through further optimization
|
||||
\item Expanded language support
|
||||
\end{itemize}
|
||||
|
||||
\subsection{Research Directions}
|
||||
\begin{itemize}
|
||||
\item Advanced reasoning mechanisms
|
||||
\item Self-supervised learning improvements
|
||||
\item Zero-shot generalization enhancement
|
||||
\item Continual learning capabilities
|
||||
\end{itemize}
|
||||
|
||||
\section{Conclusion}
|
||||
|
||||
\textbf{Zen-Omni} represents a significant advancement in AI democratization,
|
||||
delivering exceptional performance for multimodal text while maintaining
|
||||
unprecedented efficiency. Through innovative architecture design and careful optimization,
|
||||
the model achieves a balance between capability and sustainability that sets a new standard
|
||||
for responsible AI development.
|
||||
|
||||
\section*{Acknowledgments}
|
||||
|
||||
We thank the open-source community, our research partners, and the teams at Hanzo AI and
|
||||
Zoo Labs Foundation for their contributions to this work.
|
||||
|
||||
\bibliographystyle{plain}
|
||||
\bibliography{references}
|
||||
|
||||
\appendix
|
||||
|
||||
\section{Model Card}
|
||||
|
||||
\begin{table}[H]
|
||||
\centering
|
||||
\begin{tabular}{ll}
|
||||
\toprule
|
||||
\textbf{Field} & \textbf{Value} \\
|
||||
\midrule
|
||||
Model Name & Zen-Omni \\
|
||||
Version & 1.0.0 \\
|
||||
Release Date & September 2025 \\
|
||||
License & Apache 2.0 \\
|
||||
Repository & \href{https://huggingface.co/zenlm/zen-omni-30b-instruct}{huggingface.co/zenlm/zen-omni-30b-instruct} \\
|
||||
Documentation & \href{https://github.com/zenlm/zen}{github.com/zenlm/zen} \\
|
||||
Contact & research@hanzo.ai \\
|
||||
\bottomrule
|
||||
\end{tabular}
|
||||
\caption{Model Card Information}
|
||||
\end{table}
|
||||
|
||||
\end{document}
|
||||
Binary file not shown.
@@ -15,7 +15,7 @@
|
||||
|
||||
\title{\textbf{Zen Privacy: Federated Learning and Privacy-Preserving AI}\\
|
||||
\large Technical Report v2025.11}
|
||||
\author{Zen LM Research Team\\
|
||||
\author{Antje Worring, Zach Kelling \\ Zen LM Research Team\\
|
||||
\texttt{research@zenlm.org}}
|
||||
\date{November 2025}
|
||||
|
||||
|
||||
Binary file not shown.
@@ -16,7 +16,7 @@
|
||||
|
||||
\title{\textbf{Zen-Pro: A High-Performance Language Model for Complex Reasoning}\\
|
||||
\large Technical Report v2025.02}
|
||||
\author{Zen LM Research Team\\
|
||||
\author{Antje Worring, Zach Kelling \\ Zen LM Research Team\\
|
||||
\texttt{research@zenlm.org}}
|
||||
\date{February 2025}
|
||||
|
||||
|
||||
Binary file not shown.
@@ -17,7 +17,7 @@
|
||||
|
||||
\title{\textbf{BitDelta: Extreme Model Compression for the Zen Family}\\
|
||||
\large Technical Report v2025.10}
|
||||
\author{Zen LM Research Team\\
|
||||
\author{Antje Worring, Zach Kelling \\ Zen LM Research Team\\
|
||||
\texttt{research@zenlm.org}}
|
||||
\date{October 2025}
|
||||
|
||||
|
||||
Binary file not shown.
+1
-1
@@ -103,7 +103,7 @@
|
||||
}
|
||||
|
||||
\author{
|
||||
\textbf{Hanzo AI Research}$^{1}$ \quad \textbf{Zoo Labs Foundation}$^{2}$ \\[0.6em]
|
||||
Antje Worring, \textbf{Hanzo AI Research}$^{1}$ \quad \textbf{Zoo Labs Foundation}$^{2}$ \\[0.6em]
|
||||
$^{1}$Hanzo AI Inc. (Techstars '17) \quad $^{2}$Zoo Labs Foundation (501(c)(3)) \\[0.3em]
|
||||
\texttt{research@hanzo.ai} \quad \texttt{foundation@zoo.ngo} \\[0.3em]
|
||||
{\small \url{https://hanzo.ai/research/zen-reasoning}}
|
||||
|
||||
Binary file not shown.
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user