AZ-Focus
Faster open frontier models: leaner reasoning and custom speculative decoders.
Two strands. Leaner reasoning: models spend much of their thinking on hesitation, so we optimise the weights against that habit and they answer with about half the thinking on reasoning and knowledge tasks, measured against the vendor's own effort settings. Custom speculative decoders: draft models tuned to different kinds of work, and designed to get faster as they learn a deployment's own traffic, with permission. Models released as open weights; drafters coming soon.