Comment on: Handbook.md shows that long policy documents do not reliably govern agents
Drift is huge between any large spec and agent implementations.I've done a ton of testing and the model doesn't matter, fable or sol still miss a ton of detail and drift.I'm building http://engine.build which closes the gap and makes sure the implementation matches the spec.It's not the same as the satisfaction you get when solving complex problems with code yourself but writing clear specs and thinking through the problem is still very satisfying to me.
Comment on: Handbook.md shows that long policy documents do not reliably govern agents
As parent implies, they're testing the wrong control mechanism. Why are you using policies instead of real controls over the weights and inference pipeline?Well the answer is that VC-backed companies decided AI is not a domain expert tool for highly competent technical users, it's a magic oracle for the lowest common denominator. So you don't get any of the actually useful controls, just context engineering like that's fucking sane at all. It's like trying to program by navigating a git history. Not writing any new code, you don't have the ability to do that. No, you exclusively have the abili