Together AI Open-Sources OSCAR: An Attention-Aware 2-Bit KV Cache Quantization System for Long-Context LLM Serving

5/25/2026

·~1 min·5/25/2026·en·3

Quick Answer

Together AI has introduced OSCAR, a 2-bit KV cache quantization system that enhances long-context LLM serving.

Quick Take

Together AI has introduced OSCAR, a 2-bit KV cache quantization system that enhances long-context LLM serving. It achieves a 3.78-point reduction in BF16 accuracy gap on Qwen3-4B-Thinking-2507 and 1.42 points on Qwen3-8B, while offering an 8× reduction in KV memory and up to 3× decode speedup at 100K context length.

Key Points

OSCAR uses attention-aware covariance structures for optimized KV cache quantization.
It operates at 2.28 bits per KV element, significantly reducing memory usage.
The system provides a substantial decode speedup, improving performance for long-context tasks.
Benchmark results show notable accuracy improvements for Qwen3 models.
OSCAR is open-sourced, allowing broader access for LLM developers.

Article Excerpt

From source RSS / original summary

Together AI has released OSCAR (Offline Spectral Covariance-Aware Rotation), an INT2 KV cache quantization method for long-context LLM serving. Unlike prior rotation-based approaches that apply data-oblivious Hadamard transforms, OSCAR derives separate rotations for keys and values from attention-aware covariance structures estimated offline. At 2. 28 bits per KV element, OSCAR reduces the BF16 accuracy gap to 3. 78 points on Qwen3-4B-Thinking-2507 and 1.

42 points on Qwen3-8B, while delivering approximately 8× KV memory reduction and up to 3× decode speedup at 100K context length. The post Together AI Open-Sources OSCAR: An Attention-Aware 2-Bit KV Cache Quantization System for Long-Context LLM Serving appeared first on MarkTechPost.

Read on marktechpost.com

Want this in your inbox every morning?

Daily brief at your local 8am — bilingual EN/中文, free.

Subscribe — it's free

More from MarkTechPost

See more →

MarkTechPost·Asif Razzaq

4w ago

FeaturedOriginal

Meet Flash-KMeans: An IO-Aware, Exact K-Means That Runs Over 200× Faster Than FAISS on GPUs

AI Summary

Flash-KMeans is an open-source, IO-aware k-means implementation that operates over 200× faster than FAISS on NVIDIA H200 GPUs. It achieves 17.9× end-to-end and 33× speedup over cuML by optimizing distance calculations and updating mechanisms without approximating results. This advancement significantly enhances performance for data scientists and machine learning practitioners.

#AI Coding #GPU #Open Source