This page summarizes, for technical teams, the layers of the Zzeti Zeka Platform developed by Zzeti, Skyloop's AI spin-off: which controls a request passes through after it enters from a channel, how a model is chosen, where data comes from, what is measured, and how the platform is installed on-premise, in a closed air-gapped network, in a private cloud or in a hybrid topology. Skyloop Cloud, the authorized reseller and deployment partner in Türkiye, installs and operates it as a fixed-scope project.
Each layer can be reviewed on its own; together they are what turns a local model into an enterprise system.
The Zeka Chat console, WhatsApp, voice, Boards and the REST API with webhooks. Every channel passes through the same identity, access and guardrail layer; your own application calls the same agents through the API.
Expert and worker agents; skills, goals and personas; a visual workflow designer with conditional logic; the Actions Inbox for human approval (human-in-the-loop); a debug trace for every step. Agents bind to the gateway, not to a model — swapping the model does not change the agent.
The Context Lake ingests documents, tables and query results into collections; RAG retrieves the relevant chunks from those collections and hands them to the model as context. Query Bench runs data-warehouse and ERP queries through MCP, and the results become reusable context. Embedding models and the vector store run inside the installation too.
The model-agnostic LLM gateway (the AI Gateway in Zzeti): local open-weight models served with vLLM, Ollama or Llama.cpp and — only if you allow them — 30+ cloud providers behind one endpoint. Model Hub and GPU Servers pull and run the models; Fine-tune Studio adapts them to your data with LoRA; load balancing, semantic routing and semantic caching are applied at the gateway.
Agents reach ERP, CRM, data-warehouse and office tools through governed MCP (Model Context Protocol) connectors — Nebim V3, Logo, SAP, Microsoft 365, Google Workspace and custom systems; there is no direct database connection. The MCP Inspector and the Functions IDE let you test and publish new connectors and functions; the integration proxy decides which outbound connections exist at all.
Security Hub and the Access Wizard (role-based access, IAM policies), guardrails and PII masking, quality evaluation with AI Judge, sessions, traces, gateway traffic and latency in Monitor, per-LLM budgets and rate limits, an audit record for every access — all on your own instance.
The request arrives with a user and agent identity; the role hierarchy (for example Store → Region → HQ) decides which data, which connectors and which models may be reached.
Input policies, PII masking and secret detection run before the model; output policies run after the answer. The rules are independent of the model behind the gateway.
The gateway routes by policy: the local model by default; semantic routing by task; a semantic-cache hit never reaches a model at all; a cloud provider only when explicitly allowed and within cost and rate limits.
Chunks from RAG collections and the results of MCP tool calls are added to the context; inference runs on your GPU with vLLM or Ollama. Load balancing puts several model replicas behind one endpoint.
Every step is traced — token count, latency, sources used, SQL executed, agent decisions; cost is attributed per team, agent and model; AI Judge scores quality on a sample. The records stay on your instance.
What model-agnostic means in practice: agents, workflows and guardrails do not depend on the model behind the gateway. When a new open-weight model is released, it is pulled from Model Hub, compared in the Playground against the same test cases and put into service with a single routing rule — no agent is rewritten.
Representative screens with sample data — click to open at full size.



Monitor, AI Judge and the audit log run on your own instance. Observability is what turns a pilot into an operated system — and what a KVKK review asks to see.
| Component | Options | Note |
|---|---|---|
| Open-weight models | Qwen, Llama, Mistral, Gemma and other open-weight models on Hugging Face, including variants tuned for Turkish | Pulled by Model Hub; adapted to your data with LoRA in Fine-tune Studio |
| Inference engines | vLLM, Ollama, Llama.cpp | GPU Servers: local, remote or Kubernetes GPUs behind one endpoint |
| Cloud models (optional) | OpenAI, Anthropic, Google, AWS Bedrock, Azure and 30+ providers | Only when allowed in the gateway, with PII masking in front; never in an air-gapped installation |
| Embeddings & vector store | Embedding models and the vector store run inside the installation | RAG collections live in the Context Lake |
| Hardware | NVIDIA DGX Spark, GPU servers in your data center, a Kubernetes GPU pool, Apple Silicon as an edge node | Skyloop sizes it in the discovery phase and can supply the DGX Spark |
| Platform runtime | Kubernetes or Docker, installed with Terraform IaC | Security teams review the whole installation as code |
The difference between the topologies is the network boundary and the maintenance process, not the platform. KVKK controls are on by default in all of them.
The platform runs on Kubernetes or Docker in your data center; models run on your local GPUs; outbound connections are limited to what the integration proxy allows.
Zero external calls. Model files, updates and connectors are brought in through the maintenance process you approve; no internet is needed for licensing or telemetry.
When data must stay in the country but not necessarily in your building: the AWS Istanbul Local Zone or a private cloud tenancy in Türkiye — the same platform, the same controls.
Sensitive data stays on local models; non-sensitive tasks go through the gateway to the cloud providers you allow — with the same masking, access rules and cost limits.
Data flow step by step, the deployment-model comparison and the KVKK controls: on-premise local LLM guide →
Agents, workflows, guardrails and evaluations are bound to the gateway, not to a specific model. Local open-weight models served with vLLM or Ollama and — if you allow them — cloud providers are interchangeable behind the same endpoint; a model is replaced with a routing rule, not by rewriting agents. The same test cases in the Playground and AI Judge show whether the new model is better before it goes live.
Yes. The AI Gateway in Zzeti is an LLM gateway: a single API in front of local and cloud models that handles authentication, routing, load balancing, semantic caching, rate limits, budgets and logging. Applications and agents call the gateway; which model answers is a policy decision made inside your installation.
Traces for every session, agent step and model or MCP call; token counts and latency per model, agent and connector; cost per team, agent and model with budgets and limits; quality scores from AI Judge test cases and user feedback; fleet health, success rates and alerts in Monitor; and an audit record of who asked what and which data was read. All of it is stored on your own instance — nothing is sent out.
Everything: the gateway, local models, embeddings and the vector store, agents, MCP connectors inside your network, guardrails and the monitoring stack. Model files, platform updates and new connectors are packaged, checked and brought in through the maintenance process you approve — not over a live connection. Licensing and telemetry need no internet.
MCP (Model Context Protocol) is the open standard through which an agent calls tools and data sources in a governed way. In Zzeti, each ERP, CRM or data-warehouse system is exposed as an MCP connector with named functions and a schema — Nebim V3, Logo, SAP, Microsoft 365, Google Workspace or a custom system. The agent never opens a database connection; every call is scoped by role, logged with its trace, and write operations can require approval in the Actions Inbox.
In the Context Lake, as collections stored inside your installation together with the embeddings and the vector store. Collections can be fed from file shares and document systems, Microsoft 365 or Google Workspace through MCP, and from Query Bench results over ERP and data-warehouse queries. Access to a collection follows the same role hierarchy as everything else.
The Zzeti Zeka Platform is developed by Zzeti (Zzeti FZCO, Dubai), Skyloop's AI spin-off, and is an independent platform; it goes by 'Zzeti' globally and 'Zzeti Zeka' in Türkiye. Skyloop Cloud is the authorized reseller and deployment partner in Türkiye: it sizes the hardware, installs the platform on the customer's infrastructure, builds the integrations and operates it — as fixed-scope projects, not hourly consulting.
No — the platform installs on Kubernetes or on Docker, with Terraform, so your security team can review the installation as code. After the fixed-scope deployment project, your own team can operate it; Skyloop provides updates, model upgrades and support inside your perimeter with the access you grant, for as long as you want.
Tell us which systems will connect, which data must never leave and how many users you expect. Skyloop adapts the reference architecture to your topology and comes back with a sizing and a fixed-scope deployment project — not an hourly consulting quote.