On August 22, 2024, AI21 Labs announced two MoE models combining Mamba and Joint Attention, with a 256K context window, ExpertInt8 quantization, JSON mode, tool use, and Transformers integration.
On August 22, 2024, AI21 Labs announced the Jamba 1.5 Mini and Jamba 1.5 Large MoE models. The publication says they combine Mamba and Joint Attention and offer a 256K context window. It also mentions ExpertInt8 quantization, JSON mode, tool use, and integration with Transformers.
These details are relevant to engineers evaluating LLM inference options, but do not replace their own performance and suitability tests. To verify the scope and specifics, consult AI21 Labs’ original announcement and compare its published specifications and integration instructions with the models’ technical documentation.