Skip to content
Now opening to first teams. Join the waitlist

Product & architecture8 min read

Why We Built Torii: The Control Problem in Enterprise AI

AI use inside a company grows one team at a time. This is why we built a single control layer for access, cost, data flow and audit, including for the models a company runs itself.

Batuhan Zorbey Zengin

Published

On this page
  1. AI use that grows in scattered ways
  2. The price of sprawl: cost, access, data and audit
  3. Why existing tools aren't enough on their own
  4. What makes Torii different?
  5. Roadmap: lasting institutional memory
  6. Why Torii runs in your own environment
  7. What we'll share next

AI use that grows in scattered ways

In a company, LLM use rarely starts with a single decision. One team gets an API key for a customer support bot. Another team builds RAG over internal documents and picks a different model provider. The software team uses a coding assistant, and a few people connect their own MCP servers.

Six months later there are dozens of agents, several model providers, tools nobody has counted and an invoice nobody can fully trace.

Each team took the fastest route to its own problem, and each choice made sense at the time. Yet there is no shared control layer between the pieces.

More recently, companies have also started running open-weight models on their own GPUs to keep more control over their data and their costs. A team can fine-tune a model like Qwen on company data and connect it to its agents, which adds inference infrastructure to everything else that needs managing.

AI use without a shared layerEach team connects to models its own way. Nothing shared sits between the teams and the models they use.TEAMS & TOOLSMODELSSupport botInternal RAGCode assistantMCP serversProvider AProvider BProvider COwn GPUsQWEN · FINE-TUNEDNO SHARED CONTROL LAYER
Fig. 1Each team connects to models its own way. Nothing shared sits between the teams and the models they use.

The price of sprawl: cost, access, data and audit

Without a shared layer, four basic questions get hard to answer.

How much is each team, application or user spending? Input, output, cache reads and cache writes are priced differently, and when very small per-request amounts are rounded too early, the real spending picture can get lost at scale.

Who may use which tool? The tool an agent can technically reach and the tool its user is allowed to reach are two different things. When every application handles permissions on its own, the rules drift apart.

Where does the data go? If you don't know which data reaches which model, provider and region, it is hard to show that your data policy is being enforced.

Can you reconstruct what happened? Article 12 of the EU AI Act requires the high-risk AI systems within its scope to allow events to be recorded automatically over their lifetime. With disconnected records, tracing operations end to end and spotting changes made after the fact are both harder.

Companies that run their own models also have infrastructure to manage: GPU and VRAM sharing, KV cache, context management, isolation between models and scaling under load, each handled separately. KV cache is the memory where a model keeps the key/value values it has already computed for processed tokens, so it doesn't compute them again, and that memory is limited.

A model upgrade can bring another cost. If you fine-tuned one version of Qwen on company data, moving to the next version may take retraining or adaptation work to carry the same knowledge over.

Why existing tools aren't enough on their own

Teams that run into these problems usually reach for one of four kinds of tools, and each covers a different part of the problem.

LLM gateways and proxies put different model providers behind a common API, and some add budgets, access control and caching on top of routing. We needed this layer to cover more: what enters an agent's context, how tool use is carried out, and how inference resources are managed under the same company policies.

Agent frameworks and harnesses decide how an agent works, which steps it takes and which tools it calls, and they are the right place to build an application. Central control is a separate problem once several teams have each picked a different framework.

With workflow platforms, you define flows and run them. Moving existing agents and applications into one means rebuilding them, and you may not want to. We wanted a control layer that fits around what is already running.

Model serving tools such as vLLM and SGLang run models efficiently, yet company-wide resource planning, user permissions, cost policies and audit sit outside what a serving engine does. We built Torii to manage those pieces together, across all of these tools.

Where existing tools sit along the request pathEach category of tools covers one stretch of the path. Torii sits across all of it, under the same company policies.THE PATHEXISTING TOOLSApps & agentsAgent frameworksWorkflow platformsModel accessLLM gatewaysProxiesInferencevLLMSGLangtoriiPOLICYCOSTACCESSAUDITWHOLE PATH
Fig. 2Each category of tools covers one stretch of the path. Torii sits across all of it, under the same company policies.

What makes Torii different?

Torii has three core focuses: controlling the AI traffic that passes through it, much as an LLM gateway does; bringing the inference infrastructure where a company runs its own models under the same management; and, over time, building a lasting institutional memory from that usage.

The first two are where the product works today. Lasting institutional memory, and carrying model knowledge across versions, are on our development roadmap.

Torii is not an agent framework or a workflow editor. Applications keep running on their own frameworks, and Torii sits on the path their model and tool calls take. In most integrations, changing the connection point is enough; the application's architecture stays as it is.

1. Control of traffic and context

We designed Torii as a control layer for a company's AI use. Agents and applications stay where they are, and Torii sits on the path they use to talk to model providers or connected models.

Internally, we call this the Context Supply Chain. Every request a model receives is assembled from several parts: system instructions, conversation history, memory, parts coming from cache, search results and tool outputs. Together they form the context the model sees, and each of them can affect cost, accuracy and security.

The Context Supply ChainThe context a model sees is assembled from many parts, and each one affects cost, accuracy and security. Torii handles them on the way in.WHAT GOES INTO ONE REQUESTSystem instructionsConversation historyMemoryFrom cacheSearch resultsTool outputstoriiCONTEXT SUPPLY CHAINCost trackingCompany cachePolicy checkAudit recordModelSEES THE CONTEXT
Fig. 3The context a model sees is assembled from many parts, and each one affects cost, accuracy and security. Torii handles them on the way in.

Torii runs several jobs together on top of this flow:

  • Precise cost tracking: It calculates the cost of every request without losing very small amounts. Input, output, cache reads and cache writes are separate line items, because each one is priced and optimized differently.
  • Company-specific cache management: The cache strategy follows each company's traffic, weighing the provider's own cache mechanisms together with Torii's cache layer.
  • Central MCP and Skills management: Agents aren't sent long tool lists, descriptions and schemas. An agent states what it wants to do, and Torii picks the capability and the tool. If something is missing, Torii asks the agent for just that, runs the tool and returns the result.
  • Policy-based authorization: Access to models, tools and data is defined centrally for users and teams, rather than in each application's code. Permissions follow the person using the agent, whatever the agent itself can technically reach.
  • Audit records that make tampering visible: Every operation is recorded, and entries are linked in a hash chain, which helps detect whether a past record was altered after the fact.
  • Provider-independent operation: Torii works today with AWS Bedrock, Google Cloud Vertex AI and models on the company's own hardware, all under the same control rules. Whether a new model is available depends on support from its provider and from the Torii integration.

When making these decisions, we are especially careful not to add a second LLM call at every step. If the control layer itself created high cost and latency, it would make the problem it is trying to solve bigger.

Central MCP and Skills managementThe agent states what it wants and never receives tool lists or schemas. Torii picks the capability, asks only for what's missing, checks the user's permission and runs the tool.AgenttoriiTool · MCPSays what it wants donePicks capability & toolAsks for what's missingMissing detailsChecks user's permissionRuns the toolResultResult
Fig. 4The agent states what it wants and never receives tool lists or schemas. Torii picks the capability, asks only for what's missing, checks the user's permission and runs the tool.

2. The company's own models, the same control layer

Models a company runs on its own resources can be connected to Torii, which then manages their inference alongside external APIs.

On this side, the work is about GPU, VRAM, CPU, RAM and storage; KV cache and context management; isolation between models; and scaling.

Running many teams' models on shared hardware is harder than bringing up a single model. Workloads have to share resources without affecting each other, and the system has to keep up as the load changes.

In-house models pass through the same control layer as external providers' models, so the same access, cost and audit policies apply to both.

External and in-house models behind one control layerIn-house models pass through the same control layer as external ones, so access, cost and audit policies apply to both.Apps & agentstoriiONE SET OF RULESExternal providersBEDROCK · VERTEX AIYour own modelsYOUR HARDWAREGPU · VRAMKV cacheIsolationScaling
Fig. 5In-house models pass through the same control layer as external ones, so access, cost and audit policies apply to both.

3. Adapting to each company's usage

One of the things we noticed early while designing Torii is that the same optimization doesn't give the same result at every company. A cache strategy that saves a great deal at one company may make almost no difference to another company's usage, and the right model choice for one team may not fit another team's workload.

Traffic, repeat rates, data sensitivity and usage patterns differ from one company to the next. We are building Torii to adapt its behavior to each company's own traffic, and company-specific cache management, described above, is one place it already does that today.

Roadmap: lasting institutional memory

In Torii's next phase, we want to build an institutional memory a company keeps benefiting from, on top of managing its AI traffic as it happens. The topics below are not a list of current features; they are capabilities we are developing or planning.

Carrying knowledge across model versions without retraining

Say you have fine-tuned a version of Qwen on your own data. Moving to a new version may mean training or adapting again to carry the same domain knowledge over.

We are working on an approach that carries this learned knowledge between model versions without a full retraining process. Whether such a transfer is possible, and how well it works, depends on the model architecture and compatibility. For now, it is a research and development goal for Torii.

Enterprise RAG for internal data sources

We plan to build a RAG system that connects the company's data sources and serves that information according to company policies. The goal is for text-generating LLMs, internal agents, and models that generate images and video to reach the same institutional knowledge, each within its permissions.

Model training on the company's own infrastructure

We also want Torii to manage building training datasets from company data and training models on the company's own infrastructure. Data preparation and training would then run in an environment the organization defines.

Unit-based institutional memory

Each team's AI use builds up its own body of knowledge. As planned, records that comply with the organization's data policies and access rules would be evaluated at set intervals, turned into training datasets and used to update unit-specific models.

The aim is for the working knowledge each unit builds up over time to end up in its own models.

Path-based memory

We also want to build a path-based memory from agent and user interactions, in Torii's own memory infrastructure. Memory would be scoped to each agent and user instead of being one shared conversation history, and later interactions could reach the history they need within the defined scope and access rules.

Why Torii runs in your own environment

Torii is self-hosted: the control layer is installed on the company's own infrastructure. We made it that way for architectural reasons.

Prompts, content taken from documents, tool outputs and information about model requests can pass through the control layer. Routing all of this traffic through another SaaS environment operated on Torii's behalf would only move the data-control problem we want to solve to a new party.

Data you choose to send to an external model provider still reaches that provider. What changes is that Torii's control and audit layer stays in your own environment, and your organization decides which request goes to which external provider, under which rules.

Torii runs inside your environmentTorii's control and audit layer runs in your environment. Which requests reach an external provider, and under which rules, stays your decision.YOUR ENVIRONMENTApps & agentstoriiCONTROL · AUDITYour modelsOWN HARDWAREAudit logHASH-CHAINEDBy your rulesProviderEXTERNAL
Fig. 6Torii's control and audit layer runs in your environment. Which requests reach an external provider, and under which rules, stays your decision.

Because Torii sits in the middle, an outage would affect all the AI traffic passing through it. Stable operation under high concurrent load, failure handling and the ability to scale are therefore core requirements for the control layer.

What we'll share next

Upcoming posts will cover the problems we took on while building Torii and the technical decisions behind them, including LLM costs, cache and context management, how agents use MCP and tools, authorization, audit records, inference infrastructure and institutional memory.

For each one, we'll explain why we chose the approach we did, what it solves and where it falls short.

If you'd like to try Torii on your own infrastructure, you can join the waitlist at toriigate.ai.

Written by

Batuhan Zorbey Zengin

Clerion

Every AI request, priced, capped and logged.

Torii AI Gateway is opening to a first group of teams.

Join the waitlist