Research · Agent harnesses

The model is
half the system.

An AI agent is a model plus a harness: the software that gives it tools, context, memory and rules. We work on proven open-source harnesses, fit them to our models, and develop the method for fitting them to each business: its systems, its procedures and its policies, on-site or in a private cloud.

  1. EvaluationScored on real tasks
  2. GuardrailsPermissions, approvals, audit
  3. SkillsThe business's procedures, captured
  4. Tools & connectorsTheir systems, via MCP and APIs
  5. Context & memoryWhat the model sees, and when
  6. The modelAZ-Focus, tuned to the work

What we work on

Six layers, each fitted per business.

Off-the-shelf agents are built for a generic user with generic tools. Our working hypothesis: most of the value in a business deployment comes from fitting the harness, and the model, to the business. We build the layers and the method for fitting them.

⌁

Connectors

The business's systems, reachable by the agent through MCP servers and internal APIs: documents, CRM, ERP, ticketing, code and databases, with only the access each task needs.

◇

Skills

The business's procedures captured as reusable skills and templates: how they write a quote, triage a ticket, review a contract or close the month.

⛨

Guardrails

Permissions that match the business's policies: what the agent may read, write or run, approval steps for risky actions, and a full audit trail.

◎

Context engineering

What the model sees, and when: retrieval, memory and context budgets tuned to the business's data and to the strengths of the model.

⇄

Model–harness fit

Tool schemas, prompts and output formats tuned for our open models, and the model itself trained on the harness's own work, so the two fit each other.

▦

Measured on real work

An evaluation suite built from the business's real tasks. Every change to the model or the harness is scored before it ships, so improvements are proven, not assumed.


Built on open source

No black boxes.
No lock-in.

We build on open-source harnesses such as Hermes and pi, and on open standards such as the Model Context Protocol, rather than closed products.

  • Every line that touches the business's data can be inspected
  • Swap models, tools or vendors without starting again
  • Runs on-site or in the business's private cloud, next to the model
  • Built at Sister, Manchester, alongside our open models

Where it applies

What this looks like

The kinds of deployment we're building towards: the same foundations, shaped for very different businesses.

Legal & professional services

Contract review against the firm's playbook

The firm's clause library and negotiation positions as skills, a connector to matter management, and a rule that nothing leaves the firm.

People & policy

Answers from the staff handbook

HR policies and procedures as context, every answer cited to the policy it came from, and anything personal or sensitive routed to a person.

Software & engineering

Coding agents that know the codebase

The team's conventions, build and CI wired in, and a model trained on its repositories, running on its own GPUs.

Operations & support

Ticket triage and drafted replies

Answers grounded in the company's knowledge base and order systems, escalation rules from its team, and every reply reviewed or sent according to its policy.


The flywheel

Every task makes
the next one better.

Because the harness and the model run on the business's own hardware, what the agent does day to day can become training data, with the business's permission: the model learns the work, the draft model learns the traffic and gets faster, and the evaluation suite grows with every new kind of task.

Building agents
on open models?

We're looking for a few research partners to build and measure agents with, around real work.

Talk to us →