Text-to-Speech
ONNX
Chatterbox
onnxruntime
multilingual
voice-cloning

Chatterbox Multilingual V3 ONNX

This release contains two FP32 ONNX graph layouts converted from ResembleAI/chatterbox at revision 5bb1f6ee58e50c3b8d408bc82a6d3740c2db6e18. The selected V3 checkpoint is t3_mtl23ls_v3.safetensors; the decoder uses s3gen.pt. The upstream V3 Space snapshot is b21d9d062b4eda102f21975333276919935d2060. The exporter revision is b8b5f7f75436de240639e777dce2b7e26a305681. These files are modified ONNX exports, not original Resemble AI files.

Original model and attribution

Chatterbox and Chatterbox Multilingual V3 were created and released by Resemble AI. The authoritative model card, native PyTorch implementation, usage guidance, and safety notes are in the ResembleAI/chatterbox model repository. The original project acknowledges CosyVoice, HiFT-GAN, and Llama 3 as upstream work used by Chatterbox.

The upstream Chatterbox model is MIT licensed. Its license is reproduced as LICENSE, including the Resemble AI copyright notice. This repository contains modified ONNX exports maintained by KitsuMate and is not an official Resemble AI distribution.

Changes in this ONNX release

  • Exported the pinned Multilingual V3 checkpoint and decoder to FP32 ONNX.
  • Preserved a four-graph reference pipeline with separate token embeddings, language model, and conditional decoder stages.
  • Added an optimized five-graph pipeline with a slim speech encoder, merged embedding and last-position language-model graph, reusable flow preparation and flow-step graphs, and a separate vocoder.
  • Added explicit KV-cache inputs and outputs for device-resident autoregressive generation in compatible runtimes.
  • Added a six-graph split speed layout. The language model is rebuilt from the same FP32 weights with ONNX Runtime fused attention (GroupQueryAttention with rotary embedding), fused RMSNorm, and INT8 weight-only MatMuls. It reads token embeddings from a separate small embedding graph and updates a caller-owned fixed-size KV cache in place, which allows WebGPU graph capture. The layout also adds an FP16-compute flow step for GPUs and a vocoder whose sine-source phase uses a blocked prefix sum instead of CumSum.
  • Added optional companion dictionaries for Japanese kanji readings, Russian word stress, and Chinese word segmentation. The Japanese and Chinese paths use deterministic longest matching; the Russian table contains only unambiguous wordforms and makes no out-of-vocabulary guesses.
  • Preserved all Unicode scalar keys in the Cangjie map, including supplementary CJK characters outside the Basic Multilingual Plane.
  • Stored large tensors as adjacent ONNX external-data sidecars and supplied an exact role manifest and SHA-256 inventory.
  • Supplied Unity model-set contracts, shared download/install support, CPU and WebGPU execution configurations, and split-flow scheduling in KitsuMate's Unity packages. Those package sources are maintained separately from this model repository.
  • Omitted any reference voice. Supply only audio you have permission to use.
  • Did not include the upstream PerTh watermarking stage. Audio generated by the KitsuMate ONNX runtime must not be described as carrying that watermark.

Fixed graph layouts

Layout Graph roles and files
Four graph FP32 onnx/speech_encoder.onnx, onnx/embed_tokens.onnx, onnx/language_model.onnx, onnx/conditional_decoder_slim.onnx
Split FP32 (previous split layout) onnx/speech_encoder_slim.onnx, onnx/embedding_language_model_last.onnx, onnx/flow_prepare_slim.onnx, onnx/flow_step_slim.onnx, onnx/vocoder_slim.onnx
Split speed, WebGPU onnx/speech_encoder_slim.onnx, onnx/token_embedding.onnx, onnx/static_language_model_int8.onnx, onnx/flow_prepare_slim.onnx, onnx/flow_step_fp16.onnx, onnx/vocoder_prefix_sum.onnx
Split speed, CPU As above, with onnx/flow_step_slim.onnx

Each .onnx.data file is an external weight file required by the matching graph. Both layouts also use tokenizer.json; Cangjie5_TC.json is supplied for the Chinese tokenizer path. japanese_readings.tsv, russian_stress.tsv, and chinese_words.txt supply the optional language transforms. Their sources, transformations, and licenses are recorded in THIRD_PARTY_NOTICES.md. manifest.json records the exact role mapping, source revisions, file sizes, and SHA-256 values. SHA256SUMS covers every release artifact except the model card, manifest, and checksum file itself.

Use the separate ChatterboxModelSet and ChatterboxSplitModelSet types in ai.kitsumate.onnx.tts. Select the matching engine asset; changing engine does not convert one layout into the other. Current KitsuMate split engines use the split speed layouts; the previous split FP32 layout stays available for pinned older package versions. The tested split settings are language model guidance 0.5, six flow steps, and decoder guidance 0.7. The WebGPU and CPU speed layouts differ only in the flow-step graph: FP16 compute is faster on WebGPU and slower on CPU. The user accepted the 02-selected V3 listening output as the default on 2026-09-18. The earlier 06b-lm-mixed-last2-head comparison was also listening-approved as a compact reference, but it is not included in this FP32 release and still lacks end-to-end phone memory validation. Other precision configurations are not approved by those decisions.

Unity usage

Use the ai.kitsumate.onnx.tts package together with the ONNX Runtime backend. Create either KitsuMate/ONNX/TTS/Chatterbox Model Set for the four-graph layout or KitsuMate/ONNX/TTS/Chatterbox Split Model Set for the split speed layout. In the model-set Inspector, choose Download models..., select this repository and a pinned commit, then assign the matching Chatterbox engine asset to the existing TTS consumer. The two model-set types have separate installation identities and may be installed together. The split CPU and WebGPU engines use separate model sets because their flow-step graphs differ.

Every request needs a reference AudioClip supplied as TtsRequest.VoiceReference, unless the application assigns its own authorized default voice locally. Set TtsRequest.LanguageId to one of the language codes in the model-card metadata. KitsuMate's downloader discovers the language companions by their exact filenames and installs them with either graph layout. The tested split defaults are language-model guidance 0.5, six flow steps, decoder guidance 0.7, and seed 42.

The initial release contained only FP32 layouts. Full FP16, full INT4 LM, INT4 flow, dynamic INT8, and weight-only INT8 experiments on the original merged graphs were not retained because their measured quality or speed did not justify release. The later speed layout uses INT8 weights only in the rebuilt language model and FP16 compute only in the flow step; both were measured separately (see below). The accepted mixed last-two-layer compact experiment remains a research reference until its complete on-device memory path is measured. Existing multilingual, Turbo, and Nano repositories remain unchanged because no proposed cross-family export has cleared its own native-reference and listening gates.

Validation and limits

Split speed layout (2026-09-26, RTX 3070 Laptop GPU, ONNX Runtime 1.30.0, WebGPU EP 0.3.0, KitsuMate Unity runtime): four English and Polish prompts took 1.25, 1.20, 2.66, and 3.06 seconds (0.39-0.57 real-time factor after the first request) versus 5.2, 7.1, 15.7, and 20.8 seconds with the previous split FP32 layout. The CPU engine's long English request fell from 43 to 17 seconds. Whisper found no word errors in those outputs. The rebuilt FP32 language model matches the original export to about 1e-6 relative logit error; INT8 weights change logits by about 0.7% relative L2, and a five-seed comparison of four prompts showed no word-error increase over FP32. FP16 flow-step compute changed the decoded mel by about 0.2% relative L2 after chained steps. The prefix-sum vocoder phase is closer to a float64 reference (1e-5) than the original CumSum (6e-3). The static-cache graph needs every node on WebGPU for graph capture and a KV cache capacity that covers conditioning, text, and generated tokens (the KitsuMate runtime uses 1,024 positions). Voice-identity listening of the speed layout has not yet been recorded.

The release's nine ONNX graphs parse, and all 2,478 referenced external weight ranges resolve. The split FP32 CPU engine synthesized in the Unity Editor and in an isolated Windows IL2CPP player from file-backed, relative StreamingAssets paths. The player produced 61,440 finite samples at 24 kHz for a bounded 64-token English request. This establishes player functionality, not natural EOS stopping or release quality. Full-size local installs passed for both layouts: seven four-graph files were independently checked by size and SHA-256, and the 12-file split layout passed Unity's normal installer hash validation. A small ONNX fixture test passed the four-graph to split update and restore/complete paths, and staged path discovery passed for both model-set types. A Unity Editor round-trip test also imported a real tiny ONNX graph with nested external data and a tokenizer, then moved them back to file-backed data while preserving every reference. Interrupted binding now has explicit restore and keep-current actions in the normal download window, covered by an Editor test. Asset ownership coverage also verifies that renaming keeps an installation identity and duplicating cannot replace the original asset's files. After an upload, the exact Hub commit must be downloaded into a clean directory, checked against this inventory, and bound through both Unity model-set contracts before that commit is pinned in package suggestions or defaults. Android runtime checks and broader multilingual listening remain documented follow-up validation.

The Unity project also completes a clean ARM64 Android IL2CPP development build after these package changes. A follow-up signed ARM64 APK packaged a tiny ONNX fixture, its tokenizer, and nested external data under their exact assets/KitsuMateOnnxProbe/ paths; all three APK entries matched the source lengths and SHA-256 values. The probe APK contains a fixture rather than this full release. Runtime copying from Android's jar: StreamingAssets URL, synthesis, and memory remain device validation tasks.

The KitsuMate Unity implementation uses dictionary-based Japanese readings, Russian stress insertion, and Chinese segmentation. Japanese equal-cost reading conflicts, Russian homographs, and unknown forms are deliberately left unchanged; the Japanese and Chinese dictionaries are deterministic approximations rather than bit-identical ports of pykakasi and pkuseg. Custom-reference VAD preprocessing is still omitted. Performance reports from raw ONNX runs are not player performance claims.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for KitsuMate/chatterbox-multilingual-v3-onnx

Quantized
(33)
this model