{
  "turns": [
    {
      "role": "user",
      "content": "<context timezone=\"America/Chicago\" />\n<message id=\"376\" from=\"telegram-mg-17796\" sender=\"James Lopez\" time=\"May 29, 2026, 10:52 PM\">I saw this claim on the internet, can qwen 3.6 27B Q6 run on an rtx 5080? Is it the best choice for agentic coding on a 5080? How does it compare to sonnet, opus and gpt 5.5 very high? Please do some research and keep an eye out for the best agentic coding models on the 5080 and how they compare to the frontier models. Would ollama be sufficient or is some other way of hosting them locally more efficient or powerful?</message>"
    },
    {
      "role": "assistant",
      "content": "<message to=\"telegram-mg-17796\">The short answer is **no, Qwen 3.6 27B Q6 cannot run optimally on a single RTX 5080 (16GB VRAM).** While it will technically load, it is not the best choice for agentic coding on that card due to VRAM limitations. \n\nHere is the technical breakdown, the optimal local alternatives, and how they compare to frontier models.\n\n---\n\n### 1. The Math: Qwen 3.6 27B Q6 vs. RTX 5080 (16GB)\nThe RTX 5080 features **16GB GDDR7 VRAM**. \n* **Model Size:** A 27B parameter model at Q6 quantization (approx. 6.5 bits per weight with overhead) requires **~22GB** of space just for the weights.\n* **Context & KV Cache:** To use the model for coding, you need a decent context window (e.g., 16k–32k tokens). The KV cache alone will consume an additional **2GB to 4GB** of VRAM.\n* **The CPU Bottleneck:** Because 24GB–26GB of total space is needed, you would have to offload ~10GB of the model layers to your system RAM (CPU offloading). \n* **The Speed Penalty:** While your GPU GDDR7 VRAM runs at ~1,000 GB/s, DDR5 system RAM runs at only ~60–80 GB/s. Spilling over to system RAM will destroy your generation speeds, dropping your throughput from a snappy 60+ tokens per second (TPS) down to a painful **4 to 8 TPS**. For agentic coding, where the model must run repeatedly in loops, this is unusable.\n\n---\n\n### 2. The Local \"Sweet Spot\" for a 16GB RTX 5080\nTo get the blazing-fast speeds (>80 TPS) required for fluid agentic workflows, you want the model to fit **100% within the 16GB VRAM**, leaving a buffer for context.\n\nYour best local options on a 5080 are:\n1. **Qwen 2.5-Coder 14B (or Qwen 3.x 14B equivalent) at Q6_K or Q8_0:**\n   * At Q6_K, the model weights take **~11.4GB**, leaving **~4.5GB** of VRAM completely free for massive context and KV cache.\n   * At Q8_0, weights take **~15GB**, which is tight but runs fast for smaller context windows. \n   * *Performance:* Qwen's 14B Coder is an absolute beast. It punches way above its weight class, routinely beating older 30B+ models on coding benchmarks, and will run at **80–120 TPS** on your 5080.\n2. **DeepSeek-Coder-V2-Lite (16B MoE) at Q4_K_M or Q5_K_M:**\n   * Weights take **~10GB to 11.5GB**.\n   * It is highly effective at multi-file understanding and structure, and fits comfortably on a 16GB card.\n\n---\n\n### 3. Hosting: Is Ollama sufficient?\nOllama is incredibly convenient and handles CPU/GPU splitting automatically, but it is **not the most efficient or powerful way** to run local agents.\n\nFor maximum agentic performance, consider these alternatives:\n* **ExLlamaV2 (EXL2 format):** If you are on Windows or Linux, EXL2 is a quantization format optimized explicitly for NVIDIA GPUs. Running an EXL2 version of Qwen-Coder 14B (at 5.0 or 6.0 bpw) via a backend like **Tabby** or **Aphrodite Engine** will yield speeds **2x to 3x faster** than Ollama’s GGUF files.\n* **llama.cpp / KoboldCPP (for GGUFs):** If you stick to GGUFs, KoboldCPP offers superior control over **Smart Context Shifting** and **Prompt Caching**. Agents constantly send the same massive codebase context with minor changes. Prompt caching makes processing these repeated headers near-instantaneous. Ollama has basic caching, but llama.cpp/KoboldCPP allows you to fine-tune it to prevent severe lag during agent loops.\n* **vLLM (via WSL/Linux):** If you are running multiple parallel agents querying the same model, vLLM’s continuous batching and PagedAttention are the gold standard for throughput.\n\n---\n\n### 4. How Local Models Compare to Frontier Models (Sonnet, Opus, GPT-5.5)\nEven the best local 14B/27B models are **tactical, not strategic**, when compared to giants like Claude 3.5/3.7 Sonnet, Claude 4/Opus, or GPT-5.5.\n\n| Feature | Local 14B–27B (on 5080) | Frontier Models (Sonnet / GPT-5.5) |\n| :--- | :--- | :--- |\n| **Architectural Planning** | Poor. Good at editing single files, but struggles to plan changes across complex multi-file structures. | Excellent. Can map out complex architectural shifts and execute systematically. |\n| **Tool-Use Robustness** | Moderate. Agents rely on rigid JSON formatting and tool execution. Local models frequently hallucinate tool arguments or get stuck in loops when things go wrong. | High. Built-in system prompt adherence is incredibly robust, allowing deep, multi-step execution. |\n| **Context Window Quality** | Degrades quickly. Even if they support a 32k/128k limit, their reasoning and recall drop drastically once the context passes 12k tokens. | Superb. Near-perfect recall and reasoning depth even at 100k+ tokens. |\n| **Cost & Privacy** | **Winner:** 100% free, 100% private, local code never leaves your machine. | High API costs, potential data privacy concerns for proprietary codebases. |\n\n### Summary Recommendation\nKeep your **RTX 5080** running **Qwen 2.5-Coder 14B (Q6_K GGUF or 5.0 bpw EXL2)** for localized tasks: writing unit tests, refactoring individual functions, high-speed inline autocomplete, and private script editing. \n\nFor high-level autonomous agentic work (e.g., running `Claude Code` or deep repository-wide refactors), you will still want to routing those tasks to **Claude 3.5/3.7 Sonnet** to avoid agent loops, syntax errors, and planning failures.\n\nNano</message>"
    }
  ]
}
