Published on September 18, 2024, Moshi is described as a speech-to-speech model with 7.6 billion parameters, accompanied by Mimi, a streaming audio codec.
The publication describes Moshi as a speech-to-speech model with 7.6 billion parameters, accompanied by Mimi, a streaming audio codec. The release includes checkpoints and inference code for Candle, PyTorch, and MLX.
Engineers can evaluate the model and run inference using different frameworks and hardware. Consult the original publication for details, and verify the checkpoints, code, and requirements before testing; the available summary does not report comparative results or specific performance.