Why every task was green and the objective still missed
The most useful number your company can produce is the gap between work delivered and outcomes achieved. Almost no tool can produce it.
The quarter closes. Someone opens the objective and it is sitting at 60%. The room is quiet for a moment, and then the reasonable question arrives: how? Every project under it delivered. Nothing was late by more than a fortnight. The team worked hard and the work was good.
There are two available explanations at this point, and most organisations pick the wrong one.
The wrong one is that somebody misreported. The numbers must be broken, the scoring must be unfair, someone was generous with a percentage somewhere. This leads to an audit of the inputs, which finds nothing much, and to a quiet loss of faith in the scoring system.
The right one is that you have just produced the most valuable diagnostic your company is capable of generating, and you should be pleased rather than suspicious.
Delivery and outcome are different questions
“Did we do the work well?” and “did the work move the thing we cared about?” are separate questions with separate answers. They are correlated — badly executed work rarely moves an objective — but they are not the same question, and the correlation is much weaker than most planning processes assume.
A quarter where all the tasks scored well and the objective still missed means one of:
- The bet was wrong. The projects were the right size and well executed, but they were not the projects that would have moved this objective. This is a strategy finding and it is enormously valuable — it is cheaper to learn in one quarter than in four.
- The bet was right and too small. Everything worked and the objective needed three times the intervention. This is a resourcing finding.
- The objective moved for reasons outside the plan. A market shifted, a competitor did something, a dependency you did not control failed. This is a planning finding: your objective had an assumption in it that nobody wrote down.
Each of those three changes what you should do next quarter, and they change it in completely different directions. Confusing them for “someone misreported” means you will run the same quarter again.
Why most tools structurally cannot show you this
Here is the part that is not anybody's fault. The majority of OKR software computes an objective's score by rolling up the things beneath it. Tasks complete, key results advance, the objective's percentage rises as a weighted average of its children.
It is a reasonable-looking design and it has one fatal property:
If the objective's score is computed from the tasks, then the objective can never disagree with the tasks. The one signal you most needed has been defined out of existence.
A rollup can only tell you that work happened. It cannot tell you that work happened and did not matter, because it has no way to represent that state. Everything green at the bottom mathematically forces green at the top. The quarter where you learned the most is the exact quarter the tool is least able to describe.
The fix is a human judgment, deliberately
The alternative is to make the upper levels a judgment rather than an arithmetic consequence. Someone accountable looks at the objective and answers a different question from the one the delivery team answered.
In Axgenta that is four distinct ratings, each given by a different person:
- Did the work itself get done well? Rated by the task's reviewer, at the moment of completion.
- Did the project land a good outcome? Rated by the project's owner.
- Did this project deliver the goal? Rated by the goal's owner.
- Did this goal move the objective forward? Rated by the objective's owner.
Four questions, four people, one blend with published weights — 30% to goal achievement, 25% to project success, 25% to the work itself, 20% to project outcome. Because the top of the chain is entered by a person rather than derived from the bottom, the chain is allowed to disagree with itself. High task ratings sitting under a low objective rating is a legal, visible, meaningful state. That state is the finding.
Two design details that decide whether anyone trusts it
A judgment-based chain fails immediately if it feels arbitrary, so two things matter more than they look.
Missing is not zero. If a goal owner has not rated their level yet, that level is undefined — its weight is dropped and the remaining weights rescale to still sum to one. It is not scored as a failure. The alternative, treating absence as zero, punishes people for someone else's admin backlog and is the single fastest way to make a team stop believing a number.
You cannot rate yourself upward. Ratings you issue on your own upper-chain levels do not count toward your own score, while teammates still receive the shared result. Without that rule, every chain rating becomes a negotiation about someone's review.
The trade-off
This is more expensive than a rollup. It requires four people to form an opinion instead of a formula to run, and if your goal owners do not actually engage with their objectives, you will get a chain full of undefined levels and learn nothing.
It is also, frankly, uncomfortable. A rollup lets everyone finish the quarter agreeing that things went well. A judgment chain can produce a quarter where the team did excellent work and the strategy was wrong, and it puts that on a screen with someone's name against it. Organisations that are not ready to look at that will find reasons the tool is at fault.
But the alternative is worse than it sounds. Companies that cannot distinguish between “we executed badly” and “we executed well on the wrong thing” do not stop making that mistake. They repeat it, quarter after quarter, with everything green, and conclude that OKRs do not work for them.
They work fine. They just need to be able to tell you bad news. Which starts one level down, at whether the task ratings meant anything in the first place.
This is how Axgenta is built, not just how we think
Meetings become owned tasks, finished work is verified by a second person, and that single verification proves the objective moved. Fifteen days, full access, no card.
Read next
Management
Stop chasing people for updates
The status ping is the most expensive habit in management, and the reason it persists has nothing to do with discipline.
Accountability
The problem with letting people mark their own work done
Every task tool ships a checkbox that one person can tick alone. That checkbox is why you still hold status meetings.
