# From Context Engineering to Harness Engineering: Building Your Own Product Factory to Your Taste

With the weekend approaching, I figured I'd write another post.

Recently, I was talking with a colleague about what actually comes after agent loops, context engineering, and skills.

Two trends are starting to crystallize for me:

1.  Harness Engineering
    
2.  Software Factories within a Domain
    

At the same time, I think traditional context engineering is slowly losing some of its advantage.

Not because context is becoming unimportant, but because more and more approaches are trying to solve the problem automatically. Context can be compressed, reconstructed, persisted, or adjusted dynamically during reasoning. At the same time, models are becoming faster and cheaper.

Because of that, I think it will become increasingly less important to get everything exactly right on the first attempt. What matters much more is being able to automatically verify whether the result is correct and letting the process continue until it passes QA.

That is where Harness Engineering comes in for me.

Context management, tool selection, execution, persistence, routing, evaluation, recovery, and feedback loops are increasingly becoming parts of an explicit system surrounding the model.

DeepSeek is moving quite clearly in this direction with its Desktop Harness. Cordis, with its plugin architecture, shows how these kinds of systems can be composed in a modular way. Then there are meta-harness approaches, where one harness optimizes the actual harness itself, and with FrontierHarness we now even have dedicated evaluations for harness quality.

The next step from there, in my view, is Software Factories within a Domain.

At that point, the goal is no longer for an agent to complete a task. You build a production process that can continuously produce, verify, and improve software within a particular domain.

Instead of rediscovering libraries, services, documentation, rules, and knowledge for every new project, the factory might contain things like:

reusable tools and services standardized agent workflows up-to-date domain documentation architecture rules security requirements regulatory requirements UX standards automated tests, evals, and review loops

I've also linked an example of a Game Factory below. Games themselves are not the important use case here. For me, the example simply makes the idea of such a factory much easier to grasp.

In an enterprise environment, this could eventually look similar.

There might be a central Mobile UX knowledge base, a FINMA knowledge base, security rules, and other domain inventories. A product would then mostly declare which requirements apply to it and which of those requirements it currently satisfies.

If a regulatory requirement or security standard changes, the factory could automatically determine which products are affected and derive the necessary changes—or at least concrete work items—from that change.

I think the first genuinely useful factory templates are where we will start to feel some of the promised 10x productivity gains.

But the individual IC increasingly becomes QA and a process designer on the factory floor.

You spend less time inspecting every individual workpiece yourself and more time asking why something failed QA, which station in the production process is behaving incorrectly, and what needs to change in the process.

Once one layer becomes reliable, you can try moving one floor higher and automate the next part as well.

Push that far enough, and eventually everyone is playing CEO of their own little factory. What happens another few floors above that, somewhere up in the stars, remains to be seen.

For these factories to work, however, agents also need to become significantly more durable.

If the goal is no longer "complete this task" but instead "keep this production process running," you almost automatically end up with agents that operate for hours, days, or potentially indefinitely.

That is why I find developments such as Pi Durable interesting. Then there are Dots, Grokbot, and similar approaches if you don't want to deal with OpenClaw.

And this is also where the whole topic of Taste becomes interesting to me again.

If the agent produces more and more autonomously, functional correctness eventually stops being enough.

It increasingly needs to understand what kind of code I would accept, which abstractions I prefer, which trade-offs I usually make, and when something technically works but would still fail my personal QA.

All of the recent "Matter of Taste" and "Taste Driven" discussions sound very much like this to me.

Ideally, you are therefore not just building an assistant. You are building an assistant that increasingly captures your own engineering taste and consequently produces more output that passes your personal QA without intervention.

But that creates another problem.

The better such a system is supposed to become, the more context, repositories, internal information, access credentials, and eventually perhaps even personal working patterns you need to give it.

That makes Self-Hosting interesting again.

And because hardware is simultaneously scarce and expensive, another question becomes increasingly important: how much can you get out of the hardware you already own?

That is why I find Magnitude interesting, for example. It optimizes local inference for the available hardware and can sit underneath existing harnesses.

So my current thesis is roughly this:

Context is increasingly being automated. As a result, the focus shifts toward Harness Engineering. Harnesses evolve into domain-specific Software Factories. Those factories require durable agents, increasingly powerful automated QA, and eventually agents that understand your own taste. And the more data and access those systems require, the more interesting Self-Hosting becomes again.

If you want to go deeper, here are the links:

DeepSeek Desktop Harness

https://www.deepseek.com/en/harness/

Don't Blame the Large Language Model: How Scaffolding Evolution Shapes Coding Agent Quality

https://arxiv.org/abs/2607.03691

The Harness Problem

https://stencil.so/blog/the-harness-problem

Meta Harness

https://arxiv.org/abs/2603.28052

Cordis: A Programming Paradigm for Spatiotemporal Composability

https://arxiv.org/abs/2608.25512

https://github.com/cordiverse/cordis

FrontierHarness

https://frontierharness.org/

Cloudflare Clef Decision Models

https://blog.cloudflare.com/clef-decision-models/

Pi 1.0

https://earendil.com/posts/pi-1-0/

Pi Durable

https://earendil.com/posts/pi-durable/

Game Factory

https://www.bassimeledath.com/blog/game-factory

Context Language Models

https://arxiv.org/abs/2609.37725

Recursive Language Models

https://arxiv.org/abs/2512.24601

Magnitude

https://github.com/magnitudedev/magnitude
