NagibakaFrontend, bots, automation

Lesson

1.2 Which Models Are Suitable for Coding? A Cheat Sheet on Model Types: Instruct, Thinking, Densed, Multi-modal, MOE, MTP, MLX, Quantization

Written by a humanTranslated by LLM

I kept nagging one of my colleagues to try an LLM for his task of writing code for a fairly large feature. Eventually, he gave in and tried it, after which he complained for a long time and wondered how stupid the neural network was. I did not explain or tell him anything; I thought he would figure it out himself. It turned out that he was using the ChatGPT 4-mini model for coding, which is a very small and dumb model for tasks like that. For this reason, I decided that this article would not be unnecessary.

It is easy to get confused in the world of LLM models—they have become very numerous. Even if you use subscriptions, you still need to understand which model is intended for what and which model will be sufficient for a given task. Of course, you can use the smartest model for any task, but that is inefficient. Top-tier heavyweight models usually work 2–5 times slower than their smaller siblings. For the task of translating ordinary text, there is no point in using a top-tier model: a mid-range one will do just as well, but three times faster and significantly cheaper. It is like making someone with a PhD solve a school multiplication table.

When it comes to local models, you really need to understand the full variety, because you will have to find the best models for the hardware available to you. Here you will fully experience the difference, for example, between two 27-billion-parameter models: one trained almost entirely on other people's code, and the other a general-purpose model with built-in Vision. When writing code, the difference between them will be enormous.

What kinds of models are there?

Gemini_Generated_Image_y1x5wry1x5wry1x5.png

Models can be very different, and I want to tell you about the main things you need to know.

By purpose and operating logic (Interaction Types)

  • Instruct
    Such models do not have a reasoning block and therefore respond quickly. They do not continue the text; instead, they simply answer the user's request. This type of model is already outdated, but it is still effective for some use cases.

  • Thinking / Reasoning
    Reasoning models have a hidden thinking stage before producing an answer. Its duration may vary depending on the task and the model settings. These models use chains of reasoning (Chain-of-Thought) to solve complex logical tasks.

  • Multi-modal
    Multimodal models can process not only text, but also images, audio, and sometimes even video. For example, you can upload texts and images to the context simultaneously. Some models can not only accept different types of data, but also generate and return text, images, audio, and video.

    The screenshot shows an example of a small model Qwen 3.5 9B, capable of recognizing images, reasoning, and calling tools (Capabilities: Vision, Tool use, Reasoning).

image.png

By architecture and weight structure

  • Dense
    This is the classic neural network architecture. It is great for coding—offering deeper understanding than a MoEof the same size. In such models, every token (word) activates 100% of the model's parameters.

    Example: Qwen 3.6 27B

  • MoE (Mixture of Experts)
    Such models consist of a large number of layers called experts in different domains. For each token, the router activates only a subset of experts (for example, 3 out of 35). Activating all experts at once is impossible by design.

    The main advantage of this approach is speed. However, even though we use, for example, only three experts at a time, we still need to keep all available experts in memory.

    Example: Qwen3.6-35B-A3B - 3 active experts out of 35

By optimization methods and formats

  • MTP (Multi-Token Prediction)
    This is an architectural approach in which the model predicts not just one next token, but several at once. This approach speeds up text generation and improves the understanding of long-term context. It can increase token generation speed by 20–200%, depending on the complexity and routine nature of the task.

    Example: Qwen 3.6 27B MTP

  • Quantization
    Quantization is the process of compressing a model by reducing the precision of its numerical weights (for example, from 16-bit to 8-bit, 4-bit, or even 1.58-bit). For consumer hardware, an option such as Q4_K_M is usually chosen, as it is generally the most optimal choice. For example, Q8_0 will involve almost no compression and no quality loss, but it will require significantly more VRAM.

    A model is a set of floating-point numbers; quantization usually reduces the number of decimal places in these numbers: 0.3456783945673456 becomes 0.34567839, reducing it from 16 to 8 decimal places. For example, the original model weights were 32 GB and became 8 GB.

    Model generation speed is limited by memory bandwidth (VRAM). To generate 1 token, all of the model's weights need to be read. The larger the model, the longer it takes to read from memory.
    For example, let's take an nVidia RTX 5090 32GB with a memory bandwidth of 1792 GB/sec. . Let's say the model weighs 16 GB. This means we can very roughly estimate that a 16 GB model will have a generation speed of (1792/16=112), approximately 112 tokens/second, which is extremely comfortable for coding and faster than subscription-based models.

    Quantization makes it possible to run even very large models on home hardware, and some models even on smartphones. With Q4 quantization, the quality loss can be 5–10%.

image.png
  • MLX
    This is a specialized framework from Apple for efficiently running and training models. It provides a 20–40% increase in token generation speed. Optimized for Apple Silicon chips (M1/M2/M3/M4/M5+).
    Models in the MLX format use the Mac's shared memory (Unified Memory) at maximum speed.

image.png

Subscription-based cloud coding models from Anthropic (Claude Code) and ChatGPT (Codex)

I will not go into detail here. These top-tier models are used, among other things, to create new versions of themselves. They are versatile—they are suitable for coding tasks and any other tasks.

I want to show you a table using ChatGPT as an example; pay attention to the speed.

Comparison of models in the GPT-5.6 lineup (31.07.2026)

Model

Generation speed (average)

Cost (input / output per 1M tokens)

Key difference in tasks and intended use

Example of an ideal scenario

GPT-5.6 Sol (Flagship)

~45–70 tokens/sec (Up to 2.5× faster in Fast mode)

$5.00 / $30.00

Complex multi-level reasoning, in-depth code auditing, heavy analytics.

Architectural planning, identifying hidden vulnerabilities, comprehensive fintech analysis.

GPT-5.6 Terra (Balanced)

~100–120 tokens/sec

$2.50 / $15.00

A versatile “workhorse.” Ideal for standard business logic and text-related tasks.

Writing and editing articles, content plans, routine code refactoring.

GPT-5.6 Luna (Speed)

~181 tokens/sec (One of the fastest on the market)

$1.00 / $6.00 (80% cheaper than comparable models)

Ultra-fast background tasks, iterative calls in AI agent chains.

Real-time chatbot support, large-scale data structuring, bulk parsing.

Anthropic (Claude Code) has a comparable situation.

Local or enterprise coding models: Qwen 3.6 27B and Qwen 3.6-35B-A3B

Here's a small comparison table.

Comparison of Qwen 3.6 27B and Qwen 3.6-35B-A3B

Characteristic / Criterion

Qwen 3.6 27B (Dense)

Qwen 3.6-35B-A3B (MoE)

Architecture

Dense (All 27 billion parameters are always active)

Mixture of Experts (MoE) (35B total / only 3B active parameters)

Generation speed (Tokens per Second)

~25–40 tokens/sec (Up to 160 tok/s on high-end GPUs with MTP)

~105–135 tokens/sec (Up to 240+ t/s on high-end GPUs with MTP)

Input processing speed (Prefill Speed)

~800–1000 tokens/sec

~3000+ tokens/sec (Processes context 3 times faster)

VRAM requirements (For local deployment)

~18–28 GB (Fits comfortably in 24 GB with Q4_K_M quantization)

~24–38 GB (In FP8/Q8 mode, it requires more VRAM)

Strengths in tasks

Deep reasoning, error-free coding, complex agentic chains, strict adherence to long instructions.

Ultra-fast responses, RAG (working with knowledge bases), multi-user chats, initial data parsing.

Behavior when VRAM is insufficient

The speed drop is not as critical; some of it can be offloaded to RAM.

Speed drops sharply, if the MoE "experts" do not fit entirely in VRAM.

For local coding, Qwen is the favorite. But you can also consider Deepseek and GLM.

I won't overwhelm you with unnecessary information—this will be enough for now.