The Machine Climbs. The Investor Leaps.
What building an AI-assisted investment system taught us about the part of judgment that should not be automated.
8 min readIn June, Satya Nadella described what may become a company’s durable advantage once intelligence can be rented by the token.
Not the model. Everyone can rent a model.
The advantage is the loop around it: a firm’s workflows, private evaluations and tacit knowledge, fed back into a system that gets measurably better at that firm’s particular work. He called it a “hill-climbing machine” — human capital and token capital compounding together.1 He also said, almost in passing, that without human direction you get “compute running in circles.”
The phrase stayed with us because we had built one.
It also bothered us. And the caveat, we came to think, is where the real work is.
The valley hidden inside the metaphor
Hill-climbing is local search: look around, move to the adjacent point that scores higher, repeat. Its power is cadence. Cheap improvements, made often, compound.
But a hill-climber needs a score. It can only move toward whatever its objective calls “better.” And a purely local climber stops at the top of the hill it can see, even when a higher one stands beyond a valley.
There are technical ways around that — random restarts, exploration, simulated annealing. So the point is not that software can never cross a valley. The harder problem is deciding which valley is worth crossing when the decisive evidence belongs to a future that hasn’t happened yet.
That is where investing begins.
A business keeps delivering while its multiple falls. A capital-spending programme looks like either a land grab or a bonfire. A new competitor looks like an execution risk to some people and an existential one to others. Momentum, sentiment and consensus all point downhill. The investor has to decide whether another hill exists behind the fog.
That is not an escape from evidence. It is a commitment made at the edge of the evidence.
Larger and far better-resourced shops than ours have spent decades trying to mechanise that commitment itself. Some have done well. None has made the choice go away. Whatever else the record shows, it shows that.
We tried to automate the leap
We run a small investment shop on a system we built ourselves. We call it Kairos.
Kairos examines businesses against a fixed checklist. It assigns claims to confidence bands and dated grading events. It routes ideas into books with different horizons and risk rules. It keeps a scorecard of what the machine and the operator believed. It turns repeated mistakes into automated checks.
In other words, it is a hill-climbing machine for one investor.
Twice this summer, we asked it to find valley crossings for us.
First we screened for beaten-down companies: at least 30 per cent off their highs, selling pressure exhausted, a dated catalyst ahead. Before we could get attached to any of them, the system found that 87 per cent of the class failed our forensic accounting check and that the class trailed the market by roughly three points over the following year.2
Then we tried a subtler shape: businesses still delivering while their rating quietly declined, without a crash dramatic enough to set off an alarm. This was the pond where we expected to find better fish.
Across 627 historical cases, the median result was roughly ten per cent behind the market.2
We wanted the machine to prove us right. It didn’t.
That was useful. But it was not a decision.
The system could tell us about the class. It could not settle whether autonomous vehicles would destroy a ride-hailing network or enlarge it, or whether a hyperscaler’s capital spending would become a moat or a monument. The relevant facts did not yet exist.
So Kairos now does something less grand and more valuable. It prints the base rate, shows the distribution, and hands us a blank:
Name the fear. Name what would clear it. State what would prove you wrong. Sign and date it.
Then the human leaps — or doesn’t.
The machine can also teach cowardice
There is a danger in building a system that remembers every mistake.
If it grades only the leaps taken, the safest-looking behaviour is to take none. Over time, caution starts to pass for discipline. You end up with an exquisite machine for explaining why you never acted.
Nadella has said in interviews that when he became chief executive, Steve Ballmer told him: be bold, and be right. Either half is easy on its own. The job is both at once. A machine that only ever grades the leaps you took will, quietly and over years, help you with the second half by talking you out of the first.
We found three structural correctives.
First, grade the leaps not taken. Every serious candidate we decline goes into a shadow portfolio and is tracked like a position. If our passes beat our purchases, the system should say so plainly: you are becoming timid.
Second, print the tail as well as the median. “The median underperformed by ten per cent” matters. So do the winners in that same class, and what they looked like before their clouds cleared. In our own cloud-lane census the class median ran roughly ten per cent behind the market over the following year while its top decile ran about thirty-four per cent ahead — the tail the median hides.2 A base rate is a distribution, not a veto.
Third, budget the rebellion. Set aside a small, survivable allocation for well-specified judgments that run against the machine’s base case. An unused budget is information too. The question is not only “are we right?” It is “at what size can we survive being wrong three times?”
The goal is not fearlessness. It is symmetry: caution and boldness made equally visible, equally gradable, equally correctable. The machine’s filters are on probation. So is the human.
What the machine is actually for
We now think an investment machine has four jobs. None of them is the decision.
1. Print the odds before belief takes over. Show the base rate, the tail, and the closest historical cases as they looked at the moment they were detected — not with hindsight.
If you invest: before you fish in any pond, find out what the pond has actually returned, median and tail. If you build: put the base rate and the look-alike winners on the same screen as the candidate.
2. Keep the investor alive in the valley. Enforce position limits, preserve insurance, and answer uncertainty with sizing rather than false certainty.
If you invest: keep protection in a book nothing else can touch, paid for before it’s needed; ask of every leap what size survives being wrong three times. If you build: make the limits checks the system enforces, not settings the operator has to remember.
3. Remember honestly. Record what was believed, at what confidence, by what date — including the decision to pass — and grade expectancy rather than batting average.
If you invest: write every belief down as a dated claim before the outcome arrives, “I pass” included. If you build: put human judgment on the same ledger as machine output, and track what was declined.
4. Reserve attention for judgment. Automate whatever is routine so the scarce human resource goes to the premise the data cannot yet confirm — and notice when it isn’t being spent.
If you invest: count what you still do by hand each week; every repeat is a decision you’re not making. Count your leaps, too. If you build: the operator’s screen should shrink to read, decide, sign — with the year’s leap count on it.
The product is not a decision. It is a well-informed, survivable, honest leap.
A division of labour worth wanting
Technologists should build the loop. Nadella is right: the workflow, the evaluations and the accumulated learning may well become the real IP.
But do not reduce the human to the mechanic who keeps the loop running. Design for the moment when the evidence ends. Show uncertainty instead of laundering it into a score. Track refusals as carefully as actions. Make risk limits harder to override when conviction is loudest.
Investors do not need our stack for any of this. The four jobs fit in a notebook and a spreadsheet. The machine makes them cheaper. It does not make them optional.
The hopeful human-plus-AI story is not that the machine climbs while the human watches. It is that the machine clears the ground — measures the odds, keeps the record, keeps us alive — so the human can make the commitment no dataset can make on our behalf. Bold, and right. The machine can help with the second. Built honestly, it can also keep us honest about the first.
We have built the loop. We have begun grading ourselves in both columns. We have not proved an edge; the record is still far too young.
But at least we know what the machine is for.
And we know who has to jump.
Analysis and education, not investment advice.
1. Satya Nadella, post on X, June 2026, as reported by Redmond Magazine, “Nadella Says Enterprise AI’s Future May Depend Less on Frontier Models Than Learning Systems,” 18 June 2026.
2. These figures are our own in-house measurements over the screening universe our system covers — US large- and mid-cap equities — computed as medians (and, for the tail, the top decile) against a broad US market benchmark over the following twelve months, using the constituents as they stood at each detection date. The cloud class figures come from one census run of 627 historical cases: a median of roughly −10 per cent and a top decile of about +34 per cent, market-relative. They are not published studies. The samples are small enough that we treat them as priors to be updated, not as findings.