Applied AI Research Institute
We write down what would prove us wrong, before we start.
900 Labs applies one method to three questions: what is true, what a client should do, and what is worth buying. We publish how it works, and what it costs us.
Somebody was very confident, and it did not work.
Most people reading this have a version of the same story. A pilot that stopped being mentioned in the update. Nine months for something quoted at six weeks. A slide of numbers that nobody could trace back to a source when it finally mattered. A firm that was certain, and wrong, and paid either way.
Almost none of that is dishonesty. It is what happens when the people doing the analysis earn more if it comes out a particular way, and nothing in the process is designed to stop them.
So this firm starts from the other end. Before we look at anything, we write down what we expect to find and what would prove us wrong. Someone is paid to attack the answer. Someone with nothing to gain checks the numbers that carry the weight. And there is a point in every piece of work where stopping is a real option rather than an embarrassment.
We do the same thing in three places. The question changes. The reason to be careful does not.
Research
What is actually true here?
Published research with its error rate on the cover. Two or three Reports a year, Notes more often.
What would pull the answer: we published a thesis and would like it to hold.
Consulting
What should this client do?
A free diagnostic, then paid work that starts with a written diagnosis and its kill switch.
What would pull the answer: the client pays more if the answer is a big project.
The fund
What should we buy, and at what price?
Control acquisitions where the diligence is run as research.
What would pull the answer: we want the deal to work.
The third one is where this matters most. Talking yourself into a deal is the best‑documented way for a first fund to destroy itself, and we would rather have the brakes fitted before we need them.
How the method works →Three applications, and one beside them.
Research
What we publish, and at what weight
Two series, labeled so you can tell them apart before reading a word. 900 Reports are preregistered, independently verified, and published with an error rate. 900 Notes are fast pieces about what we are thinking, and they state what they did not do.
A Note may raise a question. It may not settle one.
Status One report published. One Note. Monthly is the goal.
Read the research →Consulting
A diagnosis that states what would prove it wrong
Before we start, you get a page that says what we think is wrong, what we would change, what should measurably move if we are right, and what would make us admit the diagnosis was wrong. That page is the proposal.
Sometimes it says your problem is not agentic infrastructure at all. It is that nobody owns the workflow, and that is a management fix costing nothing.
The fund
Diligence run as research
An investment thesis with walk‑away conditions written before diligence begins, a red team whose pay does not depend on the deal closing, and diligence accuracy eventually reported to investors. Control positions, 12 to 18 month holds.
Status Not raised. Vehicle not formed. Not an offer.
Read the thesis →900 Open
Software that runs on the hardware people already have
Local‑first, offline‑capable, open source, no subscription. Six repositories are public today.
See 900 Open →A diagnostic that can tell you not to buy anything.
Pick one workflow. Seven or eight questions, about ninety seconds. You get three separate readings rather than a single score, because one number cannot answer three different questions about the same process.
Opportunity
Is there enough value here to be worth pursuing?
Deployment readiness
Can your organization actually ship it?
Safe autonomy
How independently should an agent be allowed to operate?
Those three come apart more often than people expect. The most common result we designed for is a workflow that is close to perfect for automation, in an organization that should not start yet, because nobody owns the process and nobody has measured what it costs today. The tool says that plainly, then tells you which of the two to fix first.
It also refuses to flatter your arithmetic. If a claim is worth fifteen thousand euros but takes two hours to handle, the workflow is worth the two hours. Multiplying by the wrong number is how a modest process becomes a business case, and the instrument will not do it.
“I don’t know” is a valid answer to every question. It lowers the confidence of the assessment rather than the score, and sometimes it is the finding: we cannot tell whether this is worth doing, because nobody currently knows what it costs, and that is the first thing to fix.
Your result is never held behind an email address.
The diagnostic encodes a practitioner heuristic, not a validated model. Testing whether its assumptions hold is one of the purposes of our research.
In independent certification
The diagnostic is built and has been audited against sixteen adversarial cases. It goes public after independent certification, not before.
Nearly half of executives pulled back on AI agents. Half of what?
A number has been going around. Forty‑nine percent of leaders have scaled back AI agent deployments because operating costs outweighed the benefits. It usually arrives paired with a second figure: only twenty‑six percent have real‑time visibility into what their AI costs to run.
Read those two together and you form a view. They cut because they could not see.
We went and looked. The two figures come from two different surveys. One is global, with 2,145 leaders across twenty countries and territories, at companies above fifty million dollars in revenue. The other is United States only, 204 leaders, at companies above a billion, more than a third of them above ten billion. The survey carrying the twenty‑six percent contains no scale‑back figure at all.
Neither number is wrong. The sentence joining them is doing work that neither survey did.
There is also a figure in there that nobody is quoting. Organizations with full visibility into what their AI costs are five times more likely to report established returns: fifteen percent against three. Five times is the headline. The other reading of the same two numbers is that eighty‑five percent of the companies that can see their costs still have no established return.
This is a 900 Note. About two hours of work, no systematic search, no independent verification, and it says so at the top and the bottom. It raises a question. It does not settle one.
001 · The name
Where the name comes from
Stephen Hsu’s work on the genetic architecture of intelligence estimates roughly 100 standard deviations of cognitive headroom latent in the human genome, far beyond von Neumann at around six. IQ scales were never built to describe that territory. We took 900 as a name for operating in it.
Hsu, S.D.H., On the genetic architecture of intelligence and other quantitative traits, arXiv:1408.3421, section 3.3. The paper concerns human genetics rather than machine intelligence, and states no figure of 1000. We took the name, not the number.
Where we actually are
A firm that intends to publish its error rate should be able to say plainly what it has built.
- 01One published research report, The AI Execution Gap, March 2026. Thirty‑six sources, with a methodology note.
- 02A research methodology at version 2.2, revised twice after external methods review.
- 03An operating standard with a Truth Floor that is not relaxed under commercial pressure.
- 04The Agentic Readiness Diagnostic, built and audited against sixteen adversarial cases, in independent certification.
- 05Six open‑source repositories, public.
- 06A team whose credentials are real and checkable.
One thing to know before you read it
Our published report predates the current methodology. It was not preregistered and it carries no error rate. The next one will carry both. We would rather tell you than have you find out.
The credentials are real. The firm’s track record is the one thing that only time buys, and we are not going to pretend otherwise.
A note from Sam
I have spent fifteen years on both sides of this. I have paid for advice that turned out to be a decision somebody had already made, and I have been the confident person in the room who was wrong.
The thing that bothered me was never dishonesty. It was that nobody was ever surprised. The analysis kept arriving at whatever the person doing it was paid to find, and there was nothing in the process that could have stopped it.
So I built this the way I would want it done to me. We write down what would prove us wrong before we start. We pay people to attack our own conclusions. When the honest answer is that you should not spend the money, I would rather tell you that now than take the engagement.
We are new, and I am not going to dress that up. What we can show you today is how we work and what we have actually done. The rest is a record we have to earn.
Sam Rusani
Chief Executive
Start with the thing that costs you nothing.
Run the diagnostic on one workflow. If the answer is that you do not have a problem worth paying us to solve, that is what it will tell you.
Or write to a person: hello@900labs.com