Episode 145 · Enterprise · 46 min

Not everything is an LLM problem

ManageEngine's AI security lead argues the loudest model in the room is usually the wrong one: it takes anomaly detection, not an LLM, to catch the insider who pulls 490 files a day under a 500-file cap. Enterprises are late movers on generative AI for a reason — scarce data never offsets the GPU bill — and the compliance argument is that the model built for you is trained for nobody else.

SI
Sujatha Iyer
Head of AI Security, ManageEngine · with Vishal Krishna
Not everything is an LLM problem — episode thumbnail
46:27
Said in this episode
▶ 3:47
500 vs 490
Policy cap against the insider's daily pull
A file server capped at 500 downloads a day against a real baseline near 300 lets a credentialled insider take 450–490 daily without tripping a single alarm.
▶ 27:49
13+ yrs
ManageEngine's AI research history
Iyer says AI research at ManageEngine is more than thirteen years old — it was simply called big data, then Hadoop, then machine learning, deep learning, LLMs and now agentic.
▶ 38:05
85% → 94%
Where model accuracy gets hard
Open tooling reaches 80–85% accuracy readily; the climb to 93–94% requires understanding the math behind what the model is learning.
▶ 37:08
60%+
Team time spent on research
Iyer stumbles over the figure on air before settling on at least 60% of the team's time going into research — reading papers, prototyping them, and discarding them when the next one lands.
▶ 38:34
Billions/month
Calls the deployed models serve
Stated as billions of calls per month, which is why the team's work extends past model training into horizontal-versus-vertical scaling decisions.
The brief

The argument in sixty seconds

Iyer's claim is that security's move from good-to-have to must-have has been misread as a move to generative AI. Her worked example is a file server capped at 500 downloads a day: real traffic is around 300, so an insider with valid credentials pulls 490 and the rulebook never fires — the fix is a dynamic baseline learned from the data, not a bigger model. The same logic runs through the hour. Malware masquerades as a genuine executable, so detection has to read behaviour — registry keys, boot options, exfiltration — rather than signatures. Phishing emails are no longer poorly worded, because attackers hold the same LLMs everyone else does, so defence moves to domain age, IP reputation, link topology and Latin lookalike characters. But anomaly detection and forecasting are traditional machine learning, and Iyer refuses to pay the GPU tax for them: enterprises are late movers on generative AI precisely because their data is scarce, and scarce data never offsets the compute bill. LLMs earn their place in summarisation, context and the help-desk first response; agents earn theirs by stitching monitoring, ticketing and endpoint data into a single answer. Underneath sits what a CIO is actually buying: thirteen-plus years of in-house AI research, an owned data centre, no cross-organisation model sharing, PII stripped before inference, and a data protection impact assessment written for every feature. The stakes are that the last ten points of accuracy — 85 to 94 percent — are a math problem, and no amount of hype closes them.

Worth your time if you are

CIOs still policing file servers with static thresholds
CTOs being pitched an LLM for every security problem
Enterprise buyers who must ask whose data trains the model
Compliance leads mapping GDPR, HIPAA and geography
Engineering students wondering whether the math still matters
Episode map

Where the conversation travels

Every block is a chapter, coloured by what it's about. Click any of it to jump straight to that minute on YouTube.

01Cold open: good-to-have becomes must-have 0:00 Vishal frames the episode around a security-and-privacy-aware world and the posture question every new CIO faces, then hands the legacy company's dilemma — critical IT, a low-code sprint and AI on top — to ManageEngine's head of AI security. 02From static thresholds to behaviour 3:01 A decade of rule-based systems and static thresholds is undone by the insider who reads the rulebook and stays at 490 downloads under a 500 cap — and by malware that masquerades as a genuine executable, which only behavioural analysis catches. 03Digital maturity before AI maturity 6:50 AI maturity does not mushroom overnight: every interdepartmental process — the phone call to facilities, the email to payroll — has to leave a digital trail first, stitched together by a low-code layer, because AI is only as good as the data in the process. 04Masked Aadhaar, GDPR and the fines 8:56 India's passenger charts on train compartments gave way to masked Aadhaar at hotel check-in, evidence that consumer privacy expectations have hardened — and enterprises sit further ahead, where a breach costs reputation plus sector-specific fines. 05Phishing, picture-perfect 10:59 The badly spelled bonus email is gone because attackers have the same LLMs everyone else does, so detection shifts from content to features — domain registration age, IP reputation, inbound and outbound link ratios, Latin lookalike characters in a URL. 06Not everything is an LLM problem 13:55 Iyer describes Labs, the central long-term AI research team that ships capabilities into product teams' features, then draws the line she keeps returning to: spike detection and traffic forecasting are anomaly detection and old-fashioned math, not work for a language model. 07Where LLMs earn it, and the GPU tax 16:45 Generative AI belongs where summarisation, context and chain-of-thought matter — a help-desk technician's first-line response drafted from closed tickets and knowledge-base articles, or a single insight dashboard — while the enterprise's scarce data never offsets the compute cost of running large models on malware. 08What an agent actually is 20:34 Against the agentic hype that arrived on 1 January 2025, Iyer defines an agent concretely: generative and non-generative skills such as forecasting and OCR, plus contextual data stitched across monitoring and ticketing systems, plus the tools, APIs and workflows to act. 09Endpoints are not just laptops 23:18 Proactive defence means catching one infected laptop among fifty before it floods the rest and predicting the next best action — and endpoints now include the APIs everyone calls, with ransomware's real cost measured in reputation as much as ransom. 10Thirteen years, five different names 26:03 Whether you build or buy, the partner needs deep R&D rather than hype — and ManageEngine's AI is thirteen-plus years old, having been called big data, then Hadoop, then machine learning, deep learning, LLMs and now agentic, with Ask Zia embedded all along. 11Own the stack; your model is yours 28:34 Customers do not care which technique is used, only whether the job gets easier — so ManageEngine owns the data centre, applications and models in-house, forbids cross-organisation and cross-user model sharing, trains only on commercially usable datasets, anonymises PII before inference, and packages the result as Zia Agents, Agent Studio and an agent marketplace. 12Compliance checkbox, DPIA, external audits 35:01 Every enterprise evaluation opens with geography-specific compliance because the evaluator must answer to a CIO, so each feature ships with a data protection impact assessment documenting training data, scope, flow and ownership, backed by frequent internal and external audits. 13Sixty percent research, math as the floor 37:17 Most of the team's time goes to reading and prototyping papers that are superseded within days, because getting a model from 85 percent to 94 percent — and then serving billions of calls a month in production — is a math problem before it is a code problem. 14AI writes boilerplate, not judgement 40:00 Generative models can absorb boilerplate and free engineers for real knowledge work, but the developer still owns correctness, false-positive triage, retraining under data scarcity, and whether the answer is a 32GB box or a 64GB one. 15A banker's daughter and the loan rate 41:37 Iyer traces her love of math to her father, a banker who taught her what an interest rate and an MCLR were rather than calculus, and makes the case that trigonometry is visible in the pitch of a staircase and that no student is hostage to one bad teacher any more. 16Hiring for depth, and Lean In 43:22 Hiring looks for genuine depth in math or computer science and in one language rather than a list of them, and the episode closes on Sheryl Sandberg's Lean In, Michelle Obama's memoir, and Iyer's wish to do math day in and day out.
Takeaways

Ideas to carry out of this hour

01

A static threshold is a rulebook the attacker reads too

An internal file server capped at 500 downloads a day looks governed until you notice the real baseline is around 300. An insider with valid credentials — no credential stuffing, no broken authorisation, nothing to alarm on — simply sits at 450 or 490 and takes what they want. Iyer's argument is that the threshold has to be learned from the data rather than declared in a policy: a model that knows normal is 300 flags the 20 percent lift the moment it happens, which is the whole difference between a reactive and a proactive posture.

02

Detection has moved from signatures to behaviour

Malware circulating today masquerades as a genuine executable well enough that signature matching is, in her words, in for a toss. What it cannot hide is what it does — meddling with registry keys, altering boot options, staging a data exfiltration attack. The same shift has hit email: phishing is no longer poorly worded, so the features that matter are the domain's registration age, IP reputation, the ratio of inbound to outbound links, and Latin lookalike characters swapped into a URL. Attackers got the LLMs at the same time everyone else did.

03

AI maturity is downstream of a digital trail

Iyer's answer to the CIO asking where to start is not a security product. Most enterprises still run interdepartmental work by phone and email — you call facilities about the air conditioning, you email payroll about a reimbursement — and none of it leaves a record a model can read. Digitise those processes first, stitch them with a low-code layer that transcribes calls and raises tickets automatically, and security gets woven into the process rather than bolted on. AI is only as good as the data in the process.

04

Not everything is an LLM problem

Spike detection on a file server is anomaly detection. Estimating next month's traffic before a sale is forecasting. Both are traditional statistical machine learning, and Iyer is blunt that she does not want a language model doing either. Her test is per use case, not per hype cycle — LLMs and agents are genuinely useful, but the question is whether the GPU tax is worth paying, and for malware pattern analysis she says it plainly is not.

05

The enterprise is a late mover because its data is scarce

Consumer adoption of LLMs was immediate; enterprise adoption lagged, and Iyer's explanation is structural rather than cultural. Regulation and data-sharing restrictions mean enterprises have far less data to work with, and scarce data cannot offset the compute cost that large models bring. That constraint also shapes operations: with limited data you cannot simply retrain your way out of a false-positive problem, so model tuning has to be done with understanding rather than volume.

06

An agent is skills plus context plus workflow

Iyer's definition strips the word back to parts. An agent carries capabilities — generative ones, but also anomaly detection, forecasting, OCR. It carries contextual data, which in IT management means stitching server details from the monitoring system to ticket history from the help desk so a single question about one server can be answered at all. And it carries tools, APIs and workflows to act. Zia Agents ship pre-built versions, Agent Studio lets customers and partners build their own, and an agent marketplace publishes them like an app store.

07

Owning the stack is a compliance argument, not a boast

ManageEngine builds the data centre, the applications and the models in house, and Iyer connects that directly to what enterprise buyers now ask first. There is no cross-organisation or cross-user data sharing: a model built for you is only for you and will not be improved on someone else's behalf. Models, especially language models, are trained only on open-source and commercially licensed datasets; PII is stripped before data reaches a model for inference; data at rest is encrypted. Every feature carries a data protection impact assessment naming its training data, scope, flow and owner, and internal and external audits run frequently.

08

The last ten points of accuracy are a math problem

Open tooling will carry a model to 80 or 85 percent accuracy without much struggle. Getting from there to 93 or 94 requires knowing what the model is actually learning, which is why math — calculus, back-propagation — is the hiring filter and why most of the team's time goes to reading and prototyping papers that a newer paper contradicts two days later. The same discipline applies at the other end: the work does not stop at production when the system serves billions of calls a month and someone has to decide whether to scale out or up.

The numbers, drawn

What the episode measures

Every figure below was said on air — timestamps included, caveats kept.

Conversation share

portion of the hour spent on each theme
AI & machine learning · 30%SaaS & enterprise · 18%Regulation & policy · 14%Data & digitisation · 11%Product strategy · 10%Hiring & talent · 9%
AI & machine learning30%
SaaS & enterprise18%
Regulation & policy14%
Data & digitisation11%
Product strategy10%
Hiring & talent9%
Computed from the chapter map of this episode.

The rulebook the insider reads too

file downloads per day
Usual daily volume300Insider's daily pull490Policy alarm thresho500
Iyer's worked example as described on air: a static policy cap of 500 against a real baseline of roughly 300 lets a credentialled insider sit at 450–490 a day and never trip the alarm. A learned baseline would flag the 20% lift instead.▶ 3:47

Where model accuracy stops being easy

% accuracy
Open tooling gets yo85Only the math gets y94
As stated on air: 80–85% is reachable with open-source tooling, while 85% to 93–94% requires understanding what the model is learning. Upper bounds of both stated ranges shown.▶ 38:05
Worth keeping

Lines that stay

They would know the rule book — they would know that not more than 500 downloads are allowed. So they'll make 490 downloads a day, or 450, just to be sure they never cross 500. These scenarios go unnoticed.

— Sujatha Iyer ▶ 3:47

Especially in the field of security, not everything is an LLM problem. I don't want an LLM sitting and doing anomaly detection — I'd use my good old math there.

— Sujatha Iyer ▶ 15:42

If a model is built for you, it is only for you. The model will not be used for the betterment of someone else or some other organisation.

— Sujatha Iyer ▶ 31:58

Bringing a model to 80 or 85 percent accuracy is easy — there's a lot of tooling, a lot of open source. Going from 85 to 93 or 94 is difficult unless you really understand the math behind it.

— Sujatha Iyer ▶ 38:05

You could leave the code writing to AI, but you'll still have to understand what that code does. Just because the base stuff can be offloaded, it doesn't mean your foundation can be weak.

— Sujatha Iyer ▶ 40:54
Clips that travel

Short on time? Start here

CIOs still policing file servers with static thresholds

The insider who downloads 490 files

The cleanest two minutes in the episode: why a declared threshold becomes an attacker's instruction manual, and what a learned baseline catches instead.

3:01 → 5:02 · 2 min ▶ Watch clip
Security leads still training staff to spot bad spelling

Phishing is picture-perfect now

Attackers got the same LLMs everyone else did — and the features a detection engine must read once content stops being the tell.

10:59 → 13:55 · 3 min ▶ Watch clip
CTOs being pitched an LLM for every security problem

Not everything is an LLM problem

The episode's spine: anomaly detection and forecasting versus generative AI, where LLMs genuinely earn their place, and whether the GPU tax is worth paying.

14:56 → 20:34 · 6 min ▶ Watch clip
Boards worried their data is training someone else's model

Your model is trained for nobody else

No cross-user sharing, commercially licensed training data only, PII stripped before inference — with a health-plan client's ChatGPT anxiety as the live example.

30:53 → 33:45 · 3 min ▶ Watch clip
Engineers and students wondering whether math still matters

The last ten points are a math problem

Why 85 to 94 percent accuracy is where careers are made, and what a research team's week looks like when papers expire in two days.

37:17 → 38:51 · 2 min ▶ Watch clip
Glossary

The jargon, unpacked

Static threshold
A fixed, human-declared limit — no more than 500 downloads a day — that fires an alarm only when crossed, and which an insider who knows the number can simply stay beneath.
Anomaly detection
A statistical technique that learns what normal looks like for a specific signal and flags departures from it, replacing static thresholds with a baseline derived from the data.
Signature-based detection
Spotting malware by matching a known fingerprint of the file, which fails once executables can masquerade convincingly as genuine ones — hence the shift to analysing behaviour.
Data exfiltration
The unauthorised movement of data out of an organisation, whether by malware or by a credentialled insider; catching it before the data leaves is the definition of a proactive posture.
Endpoint
Any point where something touches the network — no longer just laptops and phones, but the APIs that applications call, which is why endpoint security is the weakest-link question.
DPIA
A data protection impact assessment: a per-feature document recording what data a model was trained on, the data's scope, its flow through the system, and who is accountable for it.
Agent Studio
ManageEngine's platform for customers and partners to assemble their own agents from pre-built AI capabilities, with an agent marketplace to publish them for others to use.
Connections

If this resonated, go here next

Full transcript

The whole conversation, searchable

182 segments

Auto-generated captions, lightly cleaned. Click a timestamp to open that moment on YouTube.