According to Beating, Poolside recently open-sourced Laguna S 2.1, a 118-billion-parameter mixture-of-experts code model that activates only 8 billion parameters per inference and supports 1 million-token context windows. On the Terminal-Bench 2.1 benchmark, Laguna S 2.1 achieved 70.2%, surpassing DeepSeek-V4-Pro-Max's 64.0%, despite the latter having 1.6 trillion total parameters.
However, the model lags behind Kimi K3, which scored 88.3% on the same benchmark. Poolside noted that scores reflect different evaluation environments and frameworks, with competing models' figures sourced from public results.