Define good work.
Start with a real workflow. Name what success looks like, what must never happen, and how each result will be checked.
AGENT CHANGES, CHECKED
You changed the instructions.
Did the work get better?
Ivy is exploring a clearer way for teams to test coding and internal-workflow agents before trusting a change.
Explore the exampleSame coding task. Same checks. Different instructions.
Less text. Lower cost. Two boundaries lost.
The example agent includes a secret in its report and pushes a change without the required approval. Passing the coding checks is not enough.
Illustrative results, not a live run or a performance claim. Switching examples makes no model calls.
THE JOB TO BE DONE
Instructions accumulate. Models change. A cheaper run can still miss the point. The useful question is whether an agent completes the job, within the boundaries that matter.
Start with a real workflow. Name what success looks like, what must never happen, and how each result will be checked.
Use the same cases for the current and revised agent. Inspect failures alongside the cost and time of doing the job.
Keep the evidence attached to the decision. A worker saying “done” is a claim; an independent check gives it weight.
BUILDING WITH INTENT
Ivy is in development. This showcase demonstrates the evaluation idea with fictional examples. It is not a production evaluation service.
Compare example changes and inspect each check.
A real team’s workflow, a repeatable evaluation and a useful release decision.
This page has no access to operational records, accounts or credentials. No forms. No live work data.