Skip to content
#

turboquant

Here are 134 public repositories matching this topic...

Self-hosted AI agent OS. Your memory, chat, agents, and files stay on hardware you own, offline by default, cloud by choice. Offline AI memory (taOSmd), self-hosted multi-framework group chat, a full web desktop + app store, and auto-clustering across the consumer hardware you already have (Orange/Raspberry Pi, Mac mini, gaming PC).

  • Updated Aug 2, 2026
  • Python

Native Windows vLLM 0.26.0: CPython 3.13, CUDA 12.8, SM 7.5/8.6/8.9/12.0 for RTX 20/30/40/50, OpenAI-compatible serving, Triton/FlashAttention, 10 KV-cache formats, Multi-TurboQuant, and experimental CPU/RAM/NVMe prompt-KV offload. No WSL or Docker.

  • Updated Aug 1, 2026
  • Python

Unified KV-cache compression for LLM inference: 12 Python-native methods, Debian-tested isolated add-ons, Godzilla KVarN/TriAttention, exact Godzilla/Gigatoken profiles, CUDA weight sharing, and multi-GPU planning.

  • Updated Aug 2, 2026
  • Python

Near-optimal vector quantization from Google's ICLR 2026 paper — 95% recall, 5x compression, zero preprocessing, pure Python FAISS replacement

  • Updated Mar 28, 2026
  • Python

Improve this page

Add a description, image, and links to the turboquant topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the turboquant topic, visit your repo's landing page and select "manage topics."

Learn more