tuesday, october 6, 2026 · the day's ai, attributed published by trilot llc · wyoming
archive · today in ai · 2026-08-25

IBM's Granite 4.2 adds real reasoning

Archive item — written before sources were shown.

IBM's new Granite 4.2 models add a thinking/non-thinking toggle and agentic RL training, with the 30B version beating half of SWE-Bench Verified tasks.

IBM released Granite 4.2 on August 25, its first dense, decoder-only model family built for explicit reasoning rather than direct instruction-following. Every model ships with a thinking/non-thinking toggle: switch it on for extended step-by-step reasoning before an answer, or off for a fast direct response, without swapping models.

The family comes in three sizes, 3B, 8B, and 30B parameters, all sharing the same architecture and pretraining recipe, pretrained on roughly 15 trillion tokens with a context window extended to 512K tokens. What separates the sizes is how far post-training goes: the 3B gets foundational reinforcement learning and alignment, while the 8B and 30B additionally complete an agentic RL stage, training inside real sandboxed environments where the model edits code, runs shell commands, and executes web searches rather than just predicting the next token. IBM reports meaningful results from that agentic training specifically: the 30B model scores 57.00 percent on SWE-Bench Verified and 89.17 percent on AIME25, with the 8B model at 47.67 percent and 86.67 percent respectively. All three sizes are released under the Apache 2.0 license, with weights, code, and documentation available now.

What it means for you

A reasoning toggle you control per request, rather than a separate reasoning-tuned checkpoint you have to swap to, is a genuinely practical design choice: it means one deployment can serve both fast lookups and harder multi-step problems without maintaining two model versions. The Apache 2.0 license also matters more than the benchmark numbers for most operators, since it removes the usage restrictions that come with a lot of research-grade open releases and the routing-cost tradeoffs open models often force. If you’re evaluating open models outside your usual jurisdiction or comparing against other small reasoning models for an on-device or self-hosted stack, Granite 4.2’s 30B tier is worth a real benchmark run against your own agentic workloads before you commit, not just a read of IBM’s own numbers.

sources
  1. 01Granite 4.2 LLMs: How They're Builthuggingface.co · primary, IBM Granite team
  2. 02Granite 4.2 brings native reasoning to enterprise agentsresearch.ibm.com · primary, IBM Research
Rami Steitieh
Rami Steitieh

Builder and operator. Runs 17 content sites and Trilot LLC on the tools reviewed here.