Writer's Palmyra X6 targets token costs
Archive item — written before sources were shown.
Writer's new Palmyra X6 model and upgraded agent harness aim to cut enterprise AI costs up to 50%, built on Z.ai's open-source GLM-5.2 rather than from scratch.
Enterprise AI company Writer launched Palmyra X6, a flagship model it built by post-training Z.ai’s open-source GLM-5.2 rather than training from scratch, alongside a rebuilt agentic harness aimed at cutting the cost of running both. Writer says the combined changes cut customer costs by as much as 50% on basic tasks, and cites its own research showing harness optimization alone saved an average of 40% across the multiple models it tested.
CEO May Habib framed the launch as a response to enterprise frustration with unpredictable token bills, arguing major labs still can’t deliver “flattening cost.” Writer’s pitch is that the harness, the orchestration layer that manages context, retries, and tool calls around a model, matters more to total cost than which underlying model you pick, since its efficiency gains carry over no matter which model an organization is running underneath it.
What it means for you
If you’re optimizing AI spend, Writer’s own numbers are a useful data point regardless of whether you use its product: a well-tuned harness saved 40% in its tests before any model swap. Building Palmyra X6 on top of an open-weight base model rather than training from scratch is also becoming the norm, not the exception, for vendors that want a frontier-adjacent model without frontier-lab compute budgets. Pair this with IBM’s expanded OpenAI partnership: enterprise AI vendors are converging on the same insight, that delivery and orchestration, not the model alone, is where the next round of cost and differentiation gets fought.
- 01Writer introduces new AI model and upgraded harness to contain token coststechcrunch.com · reporting
