System 1 vs LLM: Does Every AI Decision Really Need a Large Language Model?


Large Language Models have changed what we think computers can do. They can understand natural language, summarize documents, write code, analyze information, generate content and reason through surprisingly complex problems.

But as enterprises — particularly banks — begin deploying AI at scale, an interesting question emerges: Do we really need an LLM for every AI decision?

What if some decisions are much simpler? What if the information is already structured, the choices are known, the rules are clear, and all we need is a fast decision?

This question led us to experiment with an emerging idea in AI architecture: System 1 decision models. And to understand why this idea is interesting, it helps to first take a small detour into psychology.


System 1 and System 2: An Idea from Psychology

Imagine you are driving and the traffic light suddenly turns red. You don’t normally stop and carefully reason: “The traffic signal has changed from green to red. According to traffic regulations, red indicates that vehicles should stop. Therefore, I should consider applying the brakes.” You simply brake.

Now imagine someone asks: Should I buy a house or continue renting? That’s very different. You may consider interest rates, your income, property prices, future plans, family needs and many other factors. You think.

These two styles of thinking are commonly described through the System 1 / System 2 framework popularized by psychologist Daniel Kahneman. System 1 describes thinking that is fast, intuitive and automatic. System 2 describes thinking that is slower, more deliberate and analytical.

This is a model of human cognition, not a literal description of how AI systems work. But it provides a useful analogy for thinking about AI architecture.


What Would System 1 Mean for AI?

Today’s general-purpose LLMs are extraordinarily capable. Give an LLM a paragraph and ask it to understand the meaning, compare alternatives, reason about the situation and produce an explanation — and that flexibility is extremely valuable.

But enterprises also make enormous numbers of much smaller decisions. For example: Which AI model should process this request? Which tool should an AI agent call? Which agent should handle this task? Which workflow should execute next? Should this request be escalated? Where should this workload run?

These decisions can sometimes have structured inputs, explicit constraints, a finite number of possible actions, a clear objective, and a need to be made repeatedly and quickly.

That raises an architectural question: Do we need a powerful general-purpose LLM to make every one of these decisions? Or could a smaller, specialized decision model handle some of them more efficiently?

For our work, we use System 1 decision model as an engineering term for this second category: a specialized model designed to make fast, bounded decisions. We are not suggesting that these models literally reproduce human System 1 cognition. The psychology gives us the analogy.

The engineering question is much more practical:

What level of intelligence does this particular decision actually require?


We Put the Idea to the Test

We decided to investigate this through a practical banking AI problem. The problem was AI model routing.

Imagine that a bank eventually has many AI models available. Some may be highly capable but expensive. Others may be smaller and faster. Some may be permitted to process confidential information. Others may only process public information. Some may operate in particular regions. Some may be good at coding, others at summarization, reasoning or extraction.

When an application sends an AI request, somebody — or something — therefore has to answer:

Which AI model should handle this request?

We built a prototype called the Smart Model Router to study this problem. Then we compared different ways of making that decision. One of those approaches used a general-purpose OpenAI LLM. Another used Typesafe Jev 1.13, which we evaluated as a specialized fast decision model — our System 1 approach.

The code run and detailed analysis is available in the Github repo – Smart Model Router for Regulated AI


First, Governance Decides What Is Allowed

There is an important architectural principle behind our experiment: the AI router does not decide the bank’s governance rules.

Before the intelligent routing decision happens, deterministic controls decide which models are permitted. For example, suppose a request contains confidential customer information. The governance layer can check things such as data classification, presence of sensitive information, permitted providers, prohibited models, regional requirements, model capabilities, context limits, and other policy constraints.

Models that violate those requirements are removed. Only then does the decision model get involved.

The principle is simple:

Governance determines what is permitted. Intelligence determines what is preferred.

This separation is particularly important in regulated environments.


What Did the System 1 Model Actually See?

Let’s simplify the experiment. Imagine a banking application submits this request:

“Summarize this confidential financial document.”

Our router first processes the request. It identifies relevant characteristics and applies the governance rules. The result might conceptually look something like this:

Task: Summarization
Data classification: Confidential
Complexity: Medium
Reasoning required: Medium
Cost preference: Balanced
Latency requirement: 1000 ms

Permitted models:
- Model A: high quality, higher cost
- Model B: good quality, lower cost
- Model C: highest quality, higher latency

Jev then has a relatively focused job: Given these requirements and these permitted choices, which model should be selected?

That’s very different from asking an LLM to write an essay, analyze a contract or conduct an open-ended conversation. The possible actions are already known. The important information is structured. And models that the bank does not permit have already been removed.


What Happens Inside the Router?

At a simplified level, the process looks like this:

Banking AI Request
        ↓
Understand Request Properties
        ↓
Apply Hard Governance Rules
        ↓
Produce Permitted Model List
        ↓
System 1 Decision Model
        ↓
Select the Preferred Model
        ↓
Validate Decision
        ↓
Route Request

The System 1 model is therefore acting more like a specialist decision-maker than a chatbot. Its job is not to generate paragraphs of language. Its job is essentially:

Input → evaluate choices → make a decision.


Making the Comparison Fair

There was another important part of our experiment. We wanted to compare the decision-making approaches, rather than accidentally giving one approach more information than another.

So both Jev and OpenAI received the same canonical structured information and the same policy-approved candidate models. Neither received the answer we expected. Neither received the preferred model from our benchmark. And, importantly, neither received the user’s raw prompt in the primary experiment.

This created a deliberately bounded routing problem.

That last point also creates an important limitation. A general-purpose LLM’s strength is understanding language. By removing the raw prompt and giving both systems structured information, our experiment intentionally focuses on structured decision-making, not the full semantic capabilities of an LLM.

That distinction matters when interpreting the results.


So What Happened?

This is where things became interesting.

For requests that actually required an external routing decision, we measured the time required by the two live approaches. On our IID benchmark, OpenAI took 1,423.75 ms, while Jev via OpenRouter took 543.80 ms. That’s approximately a 2.62× difference in observed mean routing latency.

On our template-held-out benchmark, OpenAI took 1,406.17 ms, while Jev via OpenRouter took 562.02 ms. That’s approximately a 2.50× difference.

So across the two benchmarks, the Jev-via-OpenRouter execution path showed approximately:

2.5–2.6× lower observed mean live-routing latency.


What About Cost?

We also measured the external API cost of making the routing decisions. Across the primary experiment, the OpenAI routing API cost was approximately $0.223, while the Jev-via-OpenRouter routing API cost was approximately $0.030.

That gives an observed ratio of approximately:

7.44× lower routing API cost for the Jev-via-OpenRouter path.

These are small dollar amounts because this was an experiment. The interesting part is the ratio. If an architectural component eventually makes very large numbers of decisions, the cost and latency of each individual decision can become important.

These figures should not be interpreted as universal claims that every System 1 model will always be 2.6× faster or 7.4× cheaper than every LLM. They are the measurements we observed for the specific models, providers, configuration and experiment we evaluated.


But Did the Faster Model Make Worse Decisions?

This was perhaps the most important question. Being faster and cheaper isn’t particularly useful if the decisions are substantially worse.

We therefore compared the selected routes against our project’s synthetic reference-routing objective. In our experiment, we found:

No statistically detectable difference in Synthetic Reference-Route Agreement between Jev and OpenAI on either benchmark.

There is an important scientific distinction here. This does not prove that Jev and OpenAI are equivalent. It means our experiment did not detect a difference in reference-route agreement between them for this particular structured decision problem.

And our reference route itself is synthetic. It represents the routing objective configured for our experiment. It does not measure the quality of the answers eventually generated by the downstream AI model.

Still, combined with the latency and cost measurements, the result raises an interesting architectural question.


Perhaps the Wrong Question Is: “Which LLM Should We Use?”

When organizations discuss enterprise AI architecture, much of the conversation naturally focuses on models. Which LLM should we use? Which model is smartest? Should we use a large model or a small model? Which provider performs best?

Our experiment suggests there may be another question worth asking first:

Does this decision need a general-purpose LLM at all?

For some problems, the answer will clearly be yes. Consider interpreting an ambiguous customer request, understanding a complex document, investigating a financial situation, generating an explanation, planning a multi-step task, or reasoning about a novel situation. These problems can benefit enormously from general-purpose language models.

But consider another class of problem:

Here are ten permitted actions. Here are the constraints. Here are the characteristics of the request. Choose the most appropriate action.

That’s a very different computational problem. And it may deserve a different type of intelligence.


Why This Could Matter in Banking

A bank operating AI at scale may eventually make huge numbers of AI control decisions. Before the actual AI task even begins, systems may need to decide: Which model should handle this request? Which agent should receive this task? Which tool should the agent use? Which workflow should execute? Should this case be escalated? Where should this workload run?

Many of these decisions could potentially be bounded, repetitive and high-volume. That doesn’t mean System 1 models are automatically appropriate for all of them. It means they are worth investigating.

Imagine an enterprise architecture where different kinds of intelligence have different responsibilities. Rules enforce explicit controls. Traditional ML handles stable predictive problems where suitable training data exists. Specialized System 1 models make fast, bounded decisions. General-purpose LLMs handle language-rich, ambiguous and reasoning-intensive problems.

Instead of forcing every problem through the same technology, the architecture chooses the appropriate intelligence for the shape of the problem.


System 1 Doesn’t Replace the LLM

This is perhaps the most important takeaway.

The argument isn’t:

System 1 instead of LLMs.

It’s:

System 1 alongside LLMs.

Think about a highly skilled senior banker. You wouldn’t ask that person to personally make every tiny operational decision in the organization. Their expertise is valuable precisely because you want to use it where deeper judgement is required.

Enterprise AI may evolve in a similar direction. Use powerful general-purpose intelligence where the problem genuinely requires it. Use specialized, efficient intelligence where the decision is bounded and well defined. And use deterministic controls where the decision should not be left to AI at all.


A System 1 Decision Layer?

This leads to an architecture we think is worth exploring further. We can think of it as a System 1 Decision Layer.

It could sit between enterprise applications and more expensive or sophisticated AI capabilities. Its role would be to make high-frequency, bounded decisions quickly.

For example:

Application
     ↓
Governance
     ↓
System 1 Decision Layer
     ↓
┌────────┬────────┬────────┬────────┐
Model    Agent    Tool     Workflow
Selection Selection Selection Routing
     ↓
LLM / AI System Executes the Task

The System 1 layer isn’t necessarily doing the business task itself. It’s helping decide how the AI system should execute that task.

That distinction could become increasingly important as enterprise AI architectures become more agentic and more complex.


The Bigger Lesson From Our Experiment

Our experiment started with a relatively simple question: Can we build a smarter AI model router for banking? It ended up raising a much broader question:

What level of intelligence does each enterprise AI decision actually require?

Our results don’t provide a universal answer. They provide one useful data point.

For the bounded, structured routing problem we tested, a specialized System 1 decision approach delivered substantially lower observed latency and API cost, while we detected no statistically significant difference in reference-route agreement compared with the general-purpose LLM approach.

That makes System 1 decision intelligence an area we believe is worth exploring further. Not because LLMs are becoming less important. Quite the opposite.

As LLMs and agents become more capable, enterprises may increasingly need fast, inexpensive and specialized intelligence around them to decide when, where and how those powerful models should be used.

Perhaps the future enterprise AI architecture won’t be built around one intelligence. It will be built around layers of intelligence, each matched to the decisions it is best suited to make.

And that leads to the question we are continuing to explore:

Don’t ask only, “Which AI model should we use?”

Ask, “What level of intelligence does this decision actually require?”


Discover more from Debabrata Pruseth

Subscribe to get the latest posts sent to your email.

Scroll to Top