Episode 154 · Enterprise · 59 min

No single model gets you to the outcome

Articul8 spun out of Intel and hit its first production deployment a month before ChatGPT launched, on a bet the room disagreed with: no one model gets an enterprise to its answer. Three years on the foundation layer has closed around roughly three companies that needed $5–10 billion to enter — and Arun argues the domain-specific layer above it, where 85–90% accuracy is a failing grade, is still wide open.

A
Arun
Founder & CEO, Articul8 · with Vishal Krishna
No single model gets you to the outcome — episode thumbnail
58:43
Said in this episode
▶ 28:00
<5%
Share of the market that has actually deployed
Arun's estimate of enterprises that have got past pilots and proofs of concept — which he frames as the reason for optimism, not despair.
▶ 9:13
95%+
False positives in existing industrial alerting
Flagging systems in most industrial environments cry wolf so often that skilled technicians fall back on intuition codified in a spreadsheet.
▶ 26:17
85–90% → 95%+
Consumer accuracy versus the enterprise floor
Consumer tools get documents right 85–90% of the time; most of Articul8's industries need 95% and above just to be allowed on the field, plus provenance and repeatability.
▶ 33:23
2.8M
Documents in the first problem Articul8 solved
Cited as the deliberate choice to start with a hard problem; the smallest deployment today still runs 10,000 to 20,000 documents.
▶ 3:02
~28,000
GenAI companies in the field
Arun's framing of the noise problem — the caption garbles the phrasing, but the order of magnitude is the point he is making about differentiation.
▶ 49:48
1 in 150
Hiring conversion rate
Roughly 1:140–150, far worse than it should be, because every resume now arrives tuned almost identically to the job description.
The brief

The argument in sixty seconds

Arun's claim is that the enterprise AI fight was never at the foundation layer. That layer has already closed — really only three companies, maybe a few more, each needing five to ten billion dollars to enter and none of them meaningfully differentiated. The layer above it, domain-specific models built for manufacturing, energy, oil and gas, aerospace and large B2B financial services, is wide open, and Articul8 has been building for it since before the world had the vocabulary: its first production-scale deployment went live in November 2022, a month before ChatGPT. Domain-specific, he insists, does not mean small — a model should be as large as the data supports and as large as the task demands, and no larger. What enterprises actually buy is accuracy: consumer tools land 85–90% of the time, which is a failing grade on a plant floor where a technician's wrong call stops the line and where existing alerting already throws 95%-plus false positives. Articul8's smallest deployment is 10,000 to 20,000 documents; its first problem was 2.8 million. Around that sit the harder claims — that 70–80% of the value arrives without touching a data pipeline or an application stack, that you can delegate work but never thought, and that India's 'we don't have GPUs, so we'll build small things' modesty is a self-fulfilling prophecy in a country that aimed at Mars rather than the treetop. Fewer than 5% of enterprises have actually arrived. That, he says, is the reason for hope.

Worth your time if you are

CTOs sitting on fifty plant applications and a data swamp
Founders told the foundation-model companies will eat them
Plant engineers whose alert systems cry wolf all shift
Indian deep-tech founders arguing about GPU access
Students whose resume matches the job description exactly
Episode map

Where the conversation travels

Every block is a chapter, coloured by what it's about. Click any of it to jump straight to that minute on YouTube.

01Cold open: a spin-out timed to ChatGPT 0:00 Twenty years of training, a technology nearly ready and a market that suddenly was — Arun leaves Intel with Pat Gelsinger's backing, ships a first production deployment in November 2022 a month before ChatGPT, and opens 1 January 2024 revenue-positive with customers already in production. 02Why domain-specific, not general purpose 3:00 Amid tens of thousands of GenAI companies Articul8 picks the high-barrier, data-rich, underserved industries — manufacturing, energy, oil and gas, aerospace, B2B financial services — and holds the unfashionable position that outcomes need multiple models working together, not one giant brain. 03Three languages on one auto line 6:40 The flagship use case is an advanced technician assistant on a line where English, German and Hungarian share a shift and translation propagates errors — a system that knows the acronyms, refuses to drop a digit of a serial number, and knows you are on step 14 of 200. 04Data swamps, and plugging in between 10:00 Against the clean-the-data-first orthodoxy that built the data warehousing industry, Arun argues the first 70–80% of value arrives without changing data pipelines, sources or the application stack — Articul8 slots in between rather than becoming the new system of record. 05GPUs, no data movement, right-sized models 14:00 The only hard requirement is GPUs; data moves for processing, never for storage — and domain-specific means building the largest model the data supports and the task demands, not defaulting to small language models. 06From predictive maintenance to models watching models 16:30 The older stack needed data scientists to clean, feature-engineer, deploy and babysit decaying models — which is why roughly 90% of data science projects never left the shelf — while the real infrastructure shift is around people, not just compute at the edge. 07What an agent actually is 20:24 Arun reduces agents to a trifecta — a model, a tool and the prompts that drive one to use the other — walks through a weather query as the worked example, and notes that what used to be delightful has become table stakes. 08High-schoolers with textbooks, or an engineer 23:10 Asked how a CTO buys against 'we'll just use ChatGPT', he offers the hiring analogy — a high-schooler handed a textbook versus a 20-year engineer reading the same textbook — then the four real reasons: accuracy, provenance, repeatability and scale. 09Fewer than 5% in, and one open layer 27:00 Millions burned on stalled pilots, the slowest industries leapfrogging like India skipping CDMA for GSM, and a stack where foundation models have consolidated to about three while the domain-specific layer above stays wide open. 10Delegating thought, and the faster-horses trap 29:50 You can delegate work but not thought — code that an engineer cannot explain fails the pull request — and the way through a noisy, capital-rich market is deliberately hard problems, starting with the 2.8-million-document one, while refusing to build the faster horse customers ask for. 11One round in, and the VC herd 35:40 Incubated, one round done, revenue-positive if not cash-flow positive, with the next raise coming as the company scales into LATAM, Asia and Europe — set against a funding community that tolerates 90% failure but moves as a herd. 12India's self-induced ceiling 38:20 The 'we don't have GPUs, so we'll only do small things' posture is a self-fulfilling prophecy, and the aerospace engineer's pet peeve follows: ISRO didn't aim at the treetop, and frugal engineering means ingenuity, not cutting corners. 13Cloud wars, replayed in AI 42:20 Enterprises that swore nothing would leave their walls now route models through whichever hyperscaler they already trust — and Articul8's late-added subscription option, running inside the customer's environment, is now its fastest-growing segment. 14The refinery that found the answers rude 45:20 Models carry a deeply Western value system into an education for the next generation, and a Korean refinery pilot makes it concrete — the answers were right and nobody was happy, because only accuracy had been verified, not courtesy. 15Hiring at one in a hundred and fifty 49:20 Every resume now matches the job description exactly, conversion runs around 1:140–150, India is one of the toughest markets because candidates are brand-conscious — and candidates are invited to keep their AI open during the interview, because thinking is the thing being tested. 16Read to separate hype from reality 53:50 Sebastian Raschka for intuition and Jay Alammar for visual explanation — skip the code and the maths if you must — before a closing note on Bay Area generosity and 17 trips between 2015 and 2019.
Takeaways

Ideas to carry out of this hour

01

No single model gets you to the outcome

Articul8 committed three years ago to the position that an enterprise outcome needs several models working together, with a decision layer choosing which to call when — and was told repeatedly that it was wrong, that the large model companies would simply take over. Arun's evidence that the consensus has moved: even a company that put fifteen billion dollars into one large model now says domain-specific models will have to work together. His reasoning is stubbornly plain — human experts work that way, and these fields are deeply technical.

02

Domain-specific is not a euphemism for small

The industry has quietly collapsed 'domain-specific' into 'small language model', and Arun rejects the equation. The rule he states twice: build the largest model the data actually supports, and no larger than the task demands. Some industries hold more data than the entire internet and can support models bigger than anything shipping today; aerospace manufacturing simply does not have that much data in the world, proprietary sources included. And a model asked only to read equipment output and call a fault does not need the capacity to also write poems for fifth-graders.

03

85–90% accuracy is a consumer number

Consumer tools get things right 85 to 90% of the time when you hand them documents — and in most of the industries Articul8 sells into that is not enough to be allowed on the field. The floor to even play is 95% and above, and above that you still have to explain why the decision was made: provenance, specificity, repeatability, none of which the consumer tools carry. Scale compounds the gap. A consumer product caps at 25 documents, perhaps a couple of thousand on an enterprise tier; Articul8's smallest deployment is 10,000 to 20,000, and one plant alone holds 60 to 65 years of records.

04

Most of the value lands before the data is cleaned

A generation of enterprise IT was taught to clean the data, then structure it, then build applications — the premise the entire data warehousing industry was sold on. Arun's counter-claim is that 70 to 80% of the value now arrives without changing the data pipelines, the data sources, or in most cases anything in the application stack; the messy lake, or as he prefers, the data swamp, starts paying immediately. That also dictates the architecture: Articul8 refuses to displace the system of record the way Snowflake and Databricks grew, because a decade of hardened code will not simply be ripped out.

05

The foundation layer has closed; the layer above has not

Look at the stack, Arun argues, and the foundation model companies have already solidified into roughly three — because it takes five to ten billion dollars to play, and even then there is no differentiation left between them. The domain-specific layer above is wide open, with middleware and applications still evolving behind it. The buyers he expected to be slowest are moving fastest, because they can see the unlock and have no sunk investment to protect — the way India jumped from landlines to GSM and effectively skipped CDMA while the West hobbled between them.

06

India's ceiling is self-imposed, not compute-imposed

From his conversations, Arun diagnoses a self-induced ceiling: we don't have GPUs, so we can't build large models, so we'll aim at small things — a self-fulfilling prophecy, and one that has been endorsed by technology stalwarts. The pet peeve of an aerospace engineer follows: ISRO is rightly celebrated for reaching Mars at a fraction of the budget, but the mission was Mars, not the treetop. Frugal engineering means ingenuity, and ingenuity is not cheap. Boldness, he insists, is not a question of spend — nobody should raise $200 million to build another LLM — but of how big you are willing to think.

07

You can delegate work; you cannot delegate thought

Articul8 experiments on itself and still falls into the trap: an engineer who cannot explain a piece of code does not get the pull request through, however smart the engineer. The same test now runs through hiring, where conversion sits around 1 in 140 to 150 because every resume arrives tuned to the job description. Arun's response is to tell candidates to keep their AI assistant open during the interview and then ask questions that force them to think — if they can think with the system, that is the solution he wanted anyway. The people who introspect will separate from a crowd that has all risen to the same level.

The numbers, drawn

What the episode measures

Every figure below was said on air — timestamps included, caveats kept.

Conversation share

portion of the hour spent on each theme
AI & machine learning · 24%SaaS & enterprise · 14%Manufacturing · 12%Data & digitisation · 11%India macro · 10%Sales, GTM & growth · 9%
AI & machine learning24%
SaaS & enterprise14%
Manufacturing12%
Data & digitisation11%
India macro10%
Sales, GTM & growth9%
Computed from the chapter map of this episode.

The accuracy gap between consumer and plant floor

% correct
Consumer tools on do90Floor to even play i95Certainty on flagged98
As stated in conversation: consumer tools land 85–90% (upper bound shown), most industries need 95%+ to be allowed on the field, and Articul8 says it flags only events it is 98–99% certain about (lower bound shown).▶ 26:17

Who Articul8 hired to build a software company

% of team
Non-computer-scienti70PhDs at founding70PhDs today38
As stated in conversation: more than 70% of the founding team were non-computer-scientists and about 70% held PhDs; today PhDs are 35–40% (midpoint shown) and computer-science PhDs remain a minority.▶ 4:34
Worth keeping

Lines that stay

We are in the business of supplying engineers. We are not in the business of supplying high school students.

— Arun ▶ 25:08

You can delegate work. You cannot delegate thought — and especially in complex industries, you cannot delegate thought.

— Arun ▶ 30:27

We are in the business of inventing automobiles in a world where automobiles have never been seen. So if you go ask a customer, they're going to ask for faster horses.

— Arun ▶ 34:25

They didn't say they will reach the treetop. Their mission was to go to Mars, and they got to Mars — they figured out how to do that.

— Arun ▶ 39:23

They said: the answers are right, but they're very rude. I have no means of knowing that — all we verified was accuracy.

— Arun ▶ 48:10
Clips that travel

Short on time? Start here

Plant engineers whose alert systems cry wolf all shift

Three languages on one assembly line

The clearest picture of what a domain model actually does: shift-level multilingual records, serial numbers you cannot fumble, step 14 of 200, and 95%-plus false positives replaced by flags worth reading.

6:40 → 10:00 · 3 min ▶ Watch clip
CTOs being told the team will just use ChatGPT

High-schoolers with textbooks, or an engineer

The analogy that decides the enterprise sale, followed by the four reasons buyers actually arrive: accuracy, provenance, repeatability and document scale.

24:20 → 27:00 · 3 min ▶ Watch clip
Engineering leaders watching AI-written pull requests

Delegate the work, never the thought

The single most transferable rule in the episode, tested on Articul8's own engineers — if you cannot explain the code, it does not merge.

29:50 → 31:55 · 2 min ▶ Watch clip
Indian deep-tech founders arguing about GPU access

India's self-induced ceiling

The sharpest disagreement of the hour: why 'we lack GPUs, so we'll build small' is a self-fulfilling prophecy, and why frugal engineering was never about being cheap.

38:20 → 42:20 · 4 min ▶ Watch clip
Anyone deploying a model across languages and cultures

The refinery that found the answers rude

A deployment that passed every accuracy test and still failed, and the dialect-level nuance that no benchmark measures.

47:40 → 49:20 · 2 min ▶ Watch clip
Glossary

The jargon, unpacked

Domain-specific model
A model built for one technical field's data, vocabulary and tasks — sized to the data available and the job required, which Arun stresses is not the same thing as a small language model.
Agent
In Arun's definition, a trifecta: a model, a tool it can call, and the prompts that drive it to use the tool — plus, in bigger systems, a layer deciding which model to call when.
Foundation model
A general-purpose large model of the GPT or Claude class; Arun counts roughly three companies at that layer, each needing five to ten billion dollars to compete and little left to differentiate on.
Data lake / data swamp
A pooled store of raw enterprise data that was meant to be cleaned before use; Arun's nickname for the ones that never were, and the argument that they can start returning value anyway.
System of record
The authoritative store an enterprise runs its operations on, typically a decade of hardened code — the thing Articul8 deliberately declines to replace, plugging in between instead.
Provenance
Being able to show which source produced an answer and why the decision was made — a hard requirement in regulated and high-value industries that consumer chatbots do not meet.
Predictive maintenance
The pre-LLM industrial AI workflow — data scientists cleaning data, engineering features, deploying a model and monitoring its decay — which Arun says left roughly 90% of projects sitting on shelves.
Revenue positive
Having paying production customers from day one, which Articul8 did on 1 January 2024 — distinct, as Arun is careful to say, from being cash-flow positive.
Connections

If this resonated, go here next

Full transcript

The whole conversation, searchable

229 segments

Auto-generated captions, lightly cleaned. Click a timestamp to open that moment on YouTube.