What Insurers Need to Know About AI on the Mainframe
Phil Karecki
Field CTO
For insurance carriers, AI modernization does not always require moving data off the mainframe. This article explores how IBM Z is making it possible to run AI workloads closer to regulated systems of record, helping organizations evaluate fraud detection, risk scoring, compliance, latency, and operational efficiency as part of their modernization strategy.
Generative AI is changing how organizations think about modernization, but its impact extends beyond code and operations. It’s also influencing where workloads run and how technology architectures evolve. For insurance technology leaders, that means rethinking the relationship between AI and the systems where critical business data resides.
The April 2026 Gartner® research titled “Too Big to Fail: Why Mainframe Exit Projects Are Likely to Fail in the Age of Generative AI” identifies three separate domains where generative AI affects mainframe work: code understanding, code conversion/completion, and IT system operations. Let’s take that operational thread one step further, into architecture.
For insurance carriers running regulated systems of record, the safer long-term architecture may not be moving mainframe data to where AI happens to live. It’s bringing AI directly to where decades of that data already sits — and the newest IBM Z hardware is now purpose-built to make that practical.
The Data Gravity Problem
In our view, the Gartner report describes why this matters in large mainframe environments specifically. Because the mainframe is often the only platform capable of certain resiliency, security and transactional integrity guarantees, decades of transaction history accumulate on it, creating what the report calls “data and AI gravity.” The report notes that, “For most large-scale enterprises, the sheer volume and interconnected complexity of this data makes wholesale migration a physical and financial impossibility.”
The strategic response is already underway. There’s a shift from “moving the data to the AI, to bringing generative AI and analytical capabilities directly to the mainframe.”
For an insurance carrier, this isn’t an abstract architectural preference. Policy records, claims history and financial transaction data are exactly the kind of regulated, high-value, audit-sensitive data this argument is built around — the same heritage systems described in the second article in this series, where the risk of losing embedded institutional knowledge is precisely why wholesale replacement is the wrong instinct. Every additional hop that data takes off-platform to reach an AI model is a new question for compliance, a new latency cost, and a new place where a five-nines system of record inherits the availability and security profile of whatever it was just copied into.
The Hardware is Starting to Make this Argument Physical
Until recently, “bring AI to the mainframe” was mostly a software and integration pattern — APIs, adjacent inferencing services, careful data pipelines. That’s changing. IBM’s z17, the current generation of IBM Z hardware, is the first mainframe built with AI acceleration engineered into the processor itself rather than bolted on afterward, and it’s a useful signal of where the underlying economics of this argument are heading.
- The Telum II processor includes a second-generation on-chip AI accelerator, enabling more than 450 billion inferencing operations per day, with roughly one-millisecond response times — fast enough to score a transaction for fraud or risk in the same pass that processes it, rather than as a separate downstream step.
- The Spyre Accelerator, a PCIe-based add-on, extends this further to support generative AI and LLM workloads directly on the platform, including ensemble approaches that combine LLMs with traditional neural networks against the same enterprise data.
- AI assistants built for the platform — including IBM watsonx Code Assistant for Z and IBM watsonx Assistant for Z — extend the same code-understanding and operational-intelligence use cases described in the third article in this series, where it’s built into the environment rather than layered on top of it.
- Quantum-safe cryptography and continued resiliency improvements matter specifically for regulated industries with long data-retention horizons and active audit exposure.
The significance here isn’t any single spec. It’s that the platform vendor is now engineering for the scenario this series has been describing all along: AI as something that operates on the mainframe’s own data, under the mainframe’s own security and availability model, rather than AI being a reason to move that data somewhere else first.
A Concrete Example: Fraud Detection at the System of Record
IBM has publicly demonstrated one version of this pattern directly relevant to insurance: A home insurance claim fraud detection model that combines LLMs with neural networks in an ensemble architecture, run against the platform’s own data. The value isn’t just the model’s accuracy — it’s that the scoring happens where the claim record already lives, without a separate extract, transfer and reload cycle that widens the attack surface and adds latency between when a claim is filed and when a fraud signal is available to an adjuster.
This is a useful example precisely because it’s narrow. It doesn’t claim AI on Z solves modernization broadly, and it isn’t a substitute for the code understanding or operational intelligence work I’ve described previously. It’s one instance of the same underlying pattern applied to a workload insurance carriers will immediately recognize.
Building the System to Run it, Not Just the Model
The broader AI industry is having a version of this same realization right now, outside of mainframe specifically. After a period where enterprises rushed to stand up as many AI agents as possible, the conversation is shifting toward the harder, less exciting work of operating those agents safely, with visibility and governance, inside real business systems.
We think that shift is directly relevant here. A specialized CICS or JES agent, or a watsonx Assistant for Z deployment, is only as trustworthy as the platform discipline around it — the same resiliency, security and audit posture this series has described as a mainframe strength all along. Proximity to the system of record isn’t just a performance argument. It’s what makes an AI agent operating on regulated data governable in the first place, rather than one more integration point to secure after the fact.
Why this Matters More in Regulated Industries
For a carrier, the case for AI proximity to the system of record rests on a few factors that compound in ways they don’t for less regulated businesses:
- Data residency and audit: Claims, policy and financial transaction data often carry retention and jurisdiction requirements that get harder to satisfy the more copies and platforms that data touches.
- Latency in decisioning: Fraud scoring, underwriting risk assessment and claims triage are more valuable the closer they happen to real time — and every network hop to reach an external AI service adds to that latency.
- Security surface: Each additional integration point where regulated data leaves the mainframe’s security model is a new target and a new item on a compliance review.
- Talent and operational continuity: As described in the second and third articles in this series, the specialists who understand these systems are a shrinking population; AI tools built into the platform reduce how much of that scarce expertise has to be replicated in a separate environment.
None of this means every AI workload belongs on the mainframe, and it isn’t an argument against cloud or distributed AI, where those platforms are the better fit — that’s precisely the workload-by-workload discipline this series has argued for from the start. It means that for the specific category of AI use cases involving regulated systems of record, “move the data” should not be the default assumption it has often been.
What this Means for Insurance Executives
The practical implication is a bias worth building into modernization planning: when a use case involves regulated data on a heritage system — fraud detection, claims triage, underwriting risk scoring, real-time operational monitoring — evaluate the case for running the AI capability on the platform before assuming it needs to move. That evaluation should weigh latency, compliance exposure and total integration cost against whatever the alternative platform genuinely offers, rather than defaulting to an off-platform AI strategy because that’s where AI tooling has historically lived.
Data proximity isn’t a standalone decision, either. It’s one more variable — alongside business criticality, cost, compliance exposure and resiliency — in what’s becoming an increasingly dynamic question for every workload in the estate: not just where it runs today, but where it belongs as the platform, the regulatory environment and the available AI capabilities keep changing.
This is the fourth article in a multi-part series exploring how insurance technology leaders should think about their mainframe strategy in the age of generative AI. The first article argued that insurers are asking the wrong questions around the “mainframe exit” conversation. The second gave insurance executives a filter for telling legacy applications from heritage ones. The third mapped where generative AI genuinely helps today.
Read the Full Gartner® Report
The Gartner “Too Big to Fail: Why Mainframe Exit Projects Are Likely to Fail in the Age of Generative AI” report offers practical insights for insurance technology leaders. We feel this report is an essential read for anyone evaluating their mainframe strategy right now. Enjoy complimentary access at the link below.
Frequently Asked Questions
Why bring AI to the mainframe instead of moving mainframe data to the cloud for AI processing?
For regulated data like claims and policy records, moving data off-platform adds latency, expands the security surface, and complicates compliance and audit requirements. Running AI on the platform where the data already lives avoids all three, particularly now that IBM Z hardware includes AI acceleration built into the processor itself.
What is the Telum II processor, and why does it matter for AI on the mainframe?
Telum II is the processor at the center of IBM’s z17 mainframe. It includes a second-generation on-chip AI accelerator capable of more than 450 billion inferencing operations per day at roughly one-millisecond response times, enabling real-time scoring — such as fraud detection — in the same pass that processes a transaction.
Does bringing AI to the mainframe replace the need for cloud or distributed AI?
No. The right platform for an AI workload still depends on the workload — its data sensitivity, latency needs and regulatory exposure. AI on Z is the better fit specifically for use cases involving regulated systems of record, not a universal replacement for cloud-based AI.
How does AI proximity to the mainframe affect governance of AI agents?
Agents that operate on regulated data inherit the security, resiliency, and audit posture of whatever platform they run on. Running them close to the system of record, rather than as a separate integration point, makes that governance easier to maintain rather than something bolted on after deployment.
Gartner, Too Big to Fail: Why Mainframe Exit Projects Are Likely to Fail in the Age of Generative AI, By Dennis Smith, Alessandro Galimberti, Tobi Bet, 8 April 2026. GARTNER is a trademark of Gartner, Inc. and/or its affiliates.
Social Share
Don't miss the latest from Ensono
Keep up with Ensono
Innovation never stops, and we support you at every stage. From infrastructure-as-a-service advances to upcoming webinars, explore our news here.
Blog Post | September 15, 2026 | Inside Ensono
The 2026 IT Modernization Reality Check
Blog Post | August 27, 2026 | Inside Ensono
Where GenAI Actually Helps Mainframe Modernization
Blog Post | August 27, 2026 | Inside Ensono