Bank · Recon
Day/Night Mode

NeuralOps Integrated Enterprise AI Architecture

AI built around your real operations.

NeuralOps connects private AI infrastructure, models, agents, data intelligence, automation, and detached systems behind a locked, governed network so intelligence runs inside your business, not around it.

🏢Business App
🖥️Designated VPS
🔐Private VPN
🧠Central vLLM
🌐 Public Internet · Blocked

Direct Answer

What does AINNA AI provide?

AINNA AI provides enterprise AI infrastructure, private AI, local LLM deployment, agentic AI, automation, data intelligence and governed NeuralOps patterns for Malaysian organisations that need practical delivery, not generic demos.

NeuralOps Command Center ● SYSTEM NORMAL
Live Inference Path
App VPS VPN vLLM
Governance
Human Approval Audit Trail RBAC
Encrypted tunnel Whitelisted VPS only Public: blocked
Private AI · VPN-only inference Controlled · whitelisted VPS Sovereign · local data control Governed · human approval Detached systems · deterministic operation Agentic AI · governed execution

AI Services

From AI Infrastructure to Autonomous Business Systems

A complete service stack strategy, private infrastructure, local models, agents, automation, data intelligence, and governed deployment.

01

AI Strategy & Consulting

AI readiness assessment, use-case identification, enterprise roadmap, architecture planning, model selection, and smart routing strategy.

Discuss Deployment →
02

Private AI Infrastructure

Private vLLM, VPS infrastructure, VPN-secured access, controlled inference, private AI gateway, and model hosting with infrastructure segmentation.

View Architecture →
03

Local LLM & Model Deployment

Local model deployment, evaluation, quantized models, multi-model architecture, model routing and segmentation, inference optimization, and lifecycle management.

Explore LLM Hub →
04

AI Agents

Enterprise, operations, finance, customer-support, research, monitoring, and executive-intelligence agents plus governed multi-agent systems.

Explore Agent Hub →
05

Agentic AI Development

PLAN → DESIGN → BUILD → TEST → VALIDATE → HUMAN APPROVAL → DEPLOY → MONITOR, with audit trails and controlled execution.

Learn More →
06

AI Automation

Workflow, business-process, and data automation; API/system integration; scheduled and event-driven operations; human-approval workflows; monitoring.

Learn More →
07

Detached System Development

AI builds → validate → detach → the system runs. Deterministic systems can continue operating without continuous LLM inference where appropriate.

Learn More →
08

Data Intelligence

Data extraction, document parsing, classification, normalization, validation, reconciliation, and structured business data pipelines.

Learn More →
09

RAG & Knowledge Systems

Enterprise knowledge base, secure RAG, internal and semantic search, knowledge agents, and document retrieval with governed answers.

Learn More →
10

AI Software Development

AI-enabled web systems, business applications, internal dashboards, API and database systems, automation portals, and custom enterprise tools.

Start Pilot →
11

AI Integration

ERP, CRM, ecommerce, database, API, and legacy-system integration plus agent-to-system integration and internal systems.

View Architecture →
12

AI Governance

Human approval, RBAC, logging, audit trail, usage policy, model access control, data boundaries, monitoring, and operational guardrails.

Learn More →
13

AI Cost & Token Optimization

Smart routing, model segmentation, small-model routing, deterministic processing, detached systems, and inference optimization.

Explore LLM Strategies →
14

AI Security & Sovereignty

Private network architecture, local AI, controlled access, VPN-based inference, VPS segmentation, logging, and policy enforcement.

View Sovereignty →
15

Managed AI / AI Operations

Infrastructure, agent, workflow, model, and performance monitoring; incident detection; usage reporting; and system maintenance.

Start Pilot →

NeuralOps Service Stack

One integrated operating layer.

Each layer builds on the one below it from private infrastructure to business outcome.

Business Outcome

Operational results delivered through governed intelligence

Governance

Human approval · audit · RBAC · guardrails

Detached Systems

Deterministic systems that run after stabilization

Automation

Workflow · process · event-driven operations

AI Agents

Governed agents executing defined workflows

Data + RAG

Structured data · knowledge retrieval

Models

Local LLMs · multi-model routing

Private Infrastructure

VPS · VPN · vLLM locked network

NeuralOps Foundation

The integrated enterprise AI operating architecture

Industries

AI infrastructure built around real operations.

NeuralOps adapts to how each industry actually works with governed, private, and detached AI that fits the operational context.

🛒 Retail & Ecommerce

Inventory intelligence, marketplace operations, sales analytics, pricing, order automation, and customer-service agents.

Inventory IntelligenceSales AnalyticsPricing IntelligenceOrder AutomationDemand ForecastingMarketing AutomationDetached Ecommerce

Industry × AI Service Matrix

How services map to industries.

Select an industry to see the recommended NeuralOps services and the direct industry solution page.

Service Retail Finance Manufacturing Government Smart City
Private AI
AI Agents
Automation
Detached Systems
Data Intelligence
Governance

How Users Reach AI Without Exposing vLLM

The user never connects to the vLLM server directly. Requests go to the designated VPS first, then the VPS calls the central vLLM through a private WireGuard VPN tunnel.

1
🖥️

Designated VPS Node

Your business app, agents, dashboards, and automation run on an isolated VPS. This VPS becomes the only approved execution node for your AI workflow.

2
🔐

Private VPN Route

AINNA creates a WireGuard tunnel between your VPS and the central vLLM server. The VPS IP and tunnel credentials are whitelisted. Public traffic is rejected by design.

3
🤖

Controlled vLLM Inference

Agents send prompts through VPN, vLLM returns model output, and operations continue on the VPS. Public API keys are not required, and the model backend is not exposed in the default deployment model.

⏱ Total: 3-4 days to go live (Starter) · 5-7 days (Business/Enterprise)

AINNA NeuralOps Ecosystem Components

Complete AI infrastructure stack from local models to autonomous agents.

Local LLM Server

Private AI infrastructure running on-premise

Private LLM Infrastructure

Enterprise-grade security and data sovereignty

Agentic AI

Autonomous agents that execute complex workflows

AI Agent Builder

Design and deploy custom AI agents

Local LLM Server

Private AI infrastructure running on-premise

Private LLM Infrastructure

Enterprise-grade security and data sovereignty

Agentic AI

Autonomous agents that execute complex workflows

AI Agent Builder

Design and deploy custom AI agents

NeuralOps Service Plans

Enterprise-grade AI infrastructure with VPN-only security. Private deployment only. Designed for controlled data exposure.

Starter
RM499/month

For small businesses and startups

  • ✅ Shared VPS instance (isolated multi-tenant)
  • ✅ VPN tunnel to central vLLM server
  • ✅ Access to 2 LLM models
  • ✅ 100K tokens/day
  • ✅ Standard support (ticket/email)
  • ✅ 3-4 day onboarding
Start Pilot
Enterprise
RM3,500/month+

For enterprises and high-security needs

  • ✅ Multiple dedicated VPS instances
  • ✅ Multiple VPN tunnels (redundancy)
  • ✅ Access to ALL 7 LLM models
  • ✅ Custom rate limits
  • ✅ Full Hermes integration
  • ✅ Custom Detached System workflows
  • ✅ Dedicated account engineer
  • ✅ 99.9% uptime SLA
Contact Us

NeuralOps Service Model Super-Secured vLLM Infrastructure

Not a private LLM per customer. One centralized vLLM server. Authorized VPS instances connect via VPN. Public-facing access is disabled in the default deployment. Hub-and-spoke architecture.

LIVE Hub-and-Spoke Topology
7 Models WireGuard VPN Private Access Only
🧠
Central vLLM Server
7 LLM models · vLLM engine · GPU cluster
🔒 VPN-Only · No Public IP
WHITELISTED
🖥️
VPS Client A
Isolated · Encrypted tunnel
WHITELISTED
🖥️
VPS Client B
Isolated · Encrypted tunnel
WHITELISTED
🖥️
VPS Client N
Isolated · Encrypted tunnel
Central Inference Hub
Encrypted VPN Tunnel
Public Access Denied
🚫
No Public IP Public endpoints are disabled in the default deployment
🔐
VPN-Only Access Whitelisted VPS via encrypted tunnel
🚪
Zero Open Ports No exposed inference or admin ports
📦
VPS Isolation Each VPS isolated at cloud infra level
🔄
Snapshot Recovery Pre-change snapshots for instant rollback
🛡️
Sovereign Data Data never leaves your controlled VPS
DENIED 01

NOT Private LLM Per Customer

No per-client model isolation. Instead, a single optimized vLLM server shared via secure VPN tunnels more efficient and equally secure.

Shared Model VPN Isolated
ACTIVE 02

ONE Centralized vLLM Server

Single server running 7 LLM models. Fully utilized GPU resources. Centralized monitoring, updates, and security patches.

7 Models Central Hub
LOCKED 03
🔒

Authorized VPS Only (VPN)

Only designated VPS instances with whitelisted VPN credentials can connect. Public API keys are not required. Public endpoints are not exposed in the default deployment model.

WireGuard Whitelist
LOCKED 04
🚫

Private Connections Only

Inference ports remain private in the default deployment. Administrative access is restricted, the model backend is not publicly exposed, and direct GPU access from outside is blocked.

Private Ports No Public IP
ACTIVE 05
🌐

Hub-and-Spoke Topology

Central vLLM server connects to isolated VPS spokes via VPN. If one VPS fails, others remain unaffected. LLM server stays protected.

Hub-Spoke Fault Isolated
ACTIVE 06
📋

Audit Trail & Monitoring

Every inference request logged. Centralized monitoring across all VPS connections. Complete auditability for compliance.

PDPA Ready Full Audit
DENIED 01

NOT Private LLM Per Customer

No per-client model isolation. Instead, a single optimized vLLM server shared via secure VPN tunnels more efficient and equally secure.

Shared Model VPN Isolated
ACTIVE 02

ONE Centralized vLLM Server

Single server running 7 LLM models. Fully utilized GPU resources. Centralized monitoring, updates, and security patches.

7 Models Central Hub
LOCKED 03
🔒

Authorized VPS Only (VPN)

Only designated VPS instances with whitelisted VPN credentials can connect. Access is handled privately, and public endpoints are not exposed in the default deployment model.

WireGuard Whitelist
LOCKED 04
🚫

Private Connections Only

Inference ports remain private in the default deployment. Administrative access is restricted, the model backend is not publicly exposed, and direct GPU access from outside is blocked.

Private Ports No Public IP
ACTIVE 05
🌐

Hub-and-Spoke Topology

Central vLLM server connects to isolated VPS spokes via VPN. If one VPS fails, others remain unaffected. LLM server stays protected.

Hub-Spoke Fault Isolated
ACTIVE 06
📋

Audit Trail & Monitoring

Every inference request logged. Centralized monitoring across all VPS connections. Complete auditability for compliance.

PDPA Ready Full Audit

Data Sovereignty & Security Policy

Your data never leaves your controlled environment. All inference happens behind VPN. Third-party API calls are avoided where possible, and data exposure to public models is designed to be controlled. Security policy (kebijakan keselamatan) is enforced at the network perimeter.

📦

Your VPS

Business data stays here. Agents process, agents execute. No data leaves your isolated zone.

🔐

VPN Tunnel (Encrypted)

Only prompt text travels through. Fully encrypted. No stored logs of your data.

🧠

LLM Server

Receives prompt → returns output. No data stored. No data logged. No data shared.

Transform Raw Data Into Structured Intelligence

Supported bank-statement formats can be converted into structured financial records through dedicated parsers and validation. Sales, inventory, listings, advertising, logistics, and customer behaviour require source-specific connectors or custom parser integrations before downstream dashboards and models are built.

Data Intelligence

Extract insights from operational data

RAG Knowledge Base

Retrieval-augmented generation with your data

Vector Database

High-performance semantic search and retrieval

Detached Business System

AI-built systems that run independently

Data Intelligence

Extract insights from operational data

RAG Knowledge Base

Retrieval-augmented generation with your data

Vector Database

High-performance semantic search and retrieval

Detached Business System

AI-built systems that run independently

Intelligent Orchestration Across Your Business

AINNA automates repetitive business processes across ecommerce, inventory, reporting, listing management, operations, and internal workflows. The goal is not just automation, but intelligent orchestration between data, rules, people, and systems.

Ecommerce Automation

End-to-end online store management

Inventory Dashboard

Real-time stock tracking and alerts

Sales Report System

Automated revenue analytics and insights

Predictive Analytics

ML-powered forecasting and trends

Ecommerce Automation

End-to-end online store management

Inventory Dashboard

Real-time stock tracking and alerts

Sales Report System

Automated revenue analytics and insights

Predictive Analytics

ML-powered forecasting and trends

Bridging AI Ambition and Real-World Execution

Instead of relying only on chatbot interfaces, AINNA uses AI as a system builder designing, generating, repairing, and optimizing business applications that continue to operate independently after deployment.

🎯

Purpose-Built AI

Every system is designed for specific business outcomes, not generic conversations.

🔐

Knowledge Stays Local

Your business logic, data, and insights remain under your control no external dependencies.

♾️

Independent Operation

Systems continue running after deployment, requiring minimal maintenance and zero token costs for execution (infrastructure still applies separately).

🎯

Purpose-Built AI

Every system is designed for specific business outcomes, not generic conversations.

🔐

Knowledge Stays Local

Your business logic, data, and insights remain under your control no external dependencies.

♾️

Independent Operation

Systems continue running after deployment, requiring minimal maintenance and zero token costs for execution (infrastructure still applies separately).

AI Agents That Build Real Systems

AINNA develops AI Agents that assist in system planning, code generation, workflow design, data mapping, error detection, documentation, optimization, and system repair. These agents are not uncontrolled bots. They operate within clear business rules, human approval layers, and defined system boundaries.

🎯

System Planning & Design

Agents analyze business requirements and architect optimal system structures.

💻

Code Generation & Development

AI-assisted development produces clean, production-ready code.

🔍

Error Detection & Repair

Agents identify bugs, edge cases, and performance issues then fix them.

🛡️

Governed Operations

Every agent action follows business rules with human approval and audit trails.

🎯

System Planning & Design

Agents analyze business requirements and architect optimal system structures.

💻

Code Generation & Development

AI-assisted development produces clean, production-ready code.

🔍

Error Detection & Repair

Agents identify bugs, edge cases, and performance issues then fix them.

🛡️

Governed Operations

Every agent action follows business rules with human approval and audit trails.

Explore the AINNA Agent Hub →

Central vLLM Server Locked Behind VPN

The vLLM server is the private inference layer. It runs centrally for performance and model efficiency, but it is not visible to the public internet. Only designated VPS nodes can call it through encrypted VPN.

🔒 Private access only · Private inference ports · Restricted admin access · No direct GPU access · Whitelisted VPS only.

🏠

VPN-Secured vLLM

The vLLM server is reachable only from approved VPS nodes through private WireGuard tunnels.

🧠

Private LLM Backend

The model backend is not directly reachable from the public internet.

🎛️

Agent Execution Layer

Generic Agent AI, OpenCode, Detached Systems, and Hermes workflows run on VPS, not on the GPU server.

📈

7-Model Capacity

Central vLLM can serve multiple models while keeping access controlled through VPN-only routing.

🏠

VPN-Secured vLLM

The vLLM server is reachable only from approved VPS nodes through private WireGuard tunnels.

🧠

Private LLM Backend

The model backend is not directly reachable from the public internet.

🎛️

Agent Execution Layer

Generic Agent AI, OpenCode, Detached Systems, and Hermes workflows run on VPS, not on the GPU server.

📈

7-Model Capacity

Central vLLM can serve multiple models while keeping access controlled through VPN-only routing.

Detached Systems

Systems That Run Independently After Creation

AI designs, develops, and validates the workflow once then the detached system runs on PHP, rules, databases, and automation without repeated inference.

Explore Detached Systems →

Control-First AI Adoption

AINNA's model is control-first. Human-in-the-loop approval, audit trail, system logging, access control, role-based permissions, monitoring, observability, and business rule enforcement are included to keep AI adoption safe, explainable, and manageable.

🛡️

Governance Framework

Human-in-the-loop approval, audit trail, system logging, access control, and business rule enforcement.

🔐

Safe & Explainable AI

Every automated decision can be traced and explained with clear escalation paths.

👥

Human-in-the-Loop

Approval workflows for critical decisions with configurable risk levels.

📊

Monitoring & Audit

Complete logging, observability dashboards, and compliance reviews.

🌍

Sovereign Data & Kebijakan

Your data never leaves your controlled VPS environment. Security policy enforced at network perimeter. No third-party exposure.

🔐

Safe & Explainable AI

Every automated decision can be traced and explained with clear escalation paths.

👥

Human-in-the-Loop

Approval workflows for critical decisions with configurable risk levels.

📊

Monitoring & Audit

Complete logging, observability dashboards, and compliance reviews.

AINNA Operational Case Study

Citation-style summary based on verified operating facts from the live AINNA commerce environment.

80,000+active SKUs
9,000monthly orders
30official stores
RM15M+lifetime sales
Problem

Commerce operations required better control over catalogue breadth, order flow, store coordination, reporting and repeatable workflows across channels.

Constraints

The environment involved multiple stores, frequent product changes, operational reporting pressure and the need to keep deterministic business rules visible.

Architecture used

NeuralOps routing, detached systems, controlled validation and private AI components were used to separate repeatable work from model-heavy work.

Deterministic role

Detached systems handled routing, validation, inventory logic, order checks and other repeatable business rules before action was taken.

Outcome

The public evidence supports a practical AI operating model that fits real commerce pressure, not a purely experimental demo.

Limitations

This is an operational case study, not a controlled academic experiment. The figures should not be read as universal AI benchmark claims.

Citation information

Suggested citation: AINNA. "AINNA Operational Case Study." AINNA Research, 2026. See also Research Hub and NeuralOps Architecture.

Pitch 1: Scalable Framework for SME Transformation

Pitch 1 represents AINNA's expansion model combining business experience, AI infrastructure, detached system development, and data-driven execution into a scalable framework for SME transformation.

🎯

Phase 1: Pilot

Single use case proof of concept with measurable ROI.

📈

Phase 2: Rollout

Expand to multiple teams and integrate with core systems.

🚀

Phase 3: Scale

Organization-wide AI ecosystem with full automation.

📐

Pitch 1 Framework

Scalable framework for SME transformation.

🎯

Phase 1: Pilot

Single use case proof of concept with measurable ROI.

📈

Phase 2: Rollout

Expand to multiple teams and integrate with core systems.

🚀

Phase 3: Scale

Organization-wide AI ecosystem with full automation.

📐

Pitch 1 Framework

Scalable framework for SME transformation.

Built for Real Industries

Secure AI infrastructure tailored to your industry's compliance and data sensitivity requirements.

🛒

E-Commerce

Product description generation, customer service agents, inventory forecasting, and automated listing management across multiple stores.

⚖️

Legal Firms

Document analysis, contract review, case law research confidential data stays behind your VPN in the default deployment. No third-party exposure.

🏥

Healthcare

Patient data processing, clinical report generation, medical record analysis designed to support PDPA-aligned deployment controls, with data processing kept within the VPS in the default model.

💰

Finance

Report generation, compliance monitoring, fraud detection, risk analysis secure, auditable, with complete inference logging.

🏛️

Government

Citizen services automation, document processing, policy analysis sovereign AI infrastructure with Malaysia-based data control.

🏭

Manufacturing

IoT sensor monitoring, predictive maintenance, quality control automation, supply chain optimization with real-time agent execution.

From User to VPS to vLLM Without Public Exposure

NeuralOps protects sovereignty by separating user access, agent execution, and LLM inference. Users reach the VPS/app layer. Only the designated VPS reaches vLLM through VPN.

👤
User / Business App
Uses your app or agent UI
🌐 Public-facing layer
🖥️
Designated VPS
Agent execution + business logic
✅ Whitelisted Node
🧠
Central vLLM Server
Inference layer only
🚫 No Public Access
🌐 PUBLIC INTERNET No Access

User Stops at the VPS Layer

The user interacts with your app, dashboard, API, or agent running on VPS. The public side terminates here. The vLLM server is never exposed as a public destination.

Only Whitelisted VPS Can Enter VPN

The VPS uses WireGuard credentials and an approved IP route. If traffic does not originate from an authorized VPS tunnel, the vLLM layer rejects it.

vLLM Performs Inference Only

The LLM server receives a controlled prompt over VPN and returns output. It is not used as a storage layer, web app layer, or public integration point.

How NeuralOps Guarantees Data Sovereignty

🇲🇾

Malaysia-Based Infrastructure

All servers and VPS instances are hosted within Malaysia's borders. Your data is subject to Malaysian law (PDPA 2010), not foreign jurisdictions like US Cloud Act or GDPR.

🔒

No Third-Party API Calls

Unlike solutions that proxy through OpenAI/Google/xAI APIs, NeuralOps makes zero external API calls. Your prompt never reaches a foreign server. Inference is 100% local to our infrastructure.

📦

VPS-Level Data Isolation

Each customer's VPS is isolated at the cloud hypervisor level. No other tenant can access your files, processes, or memory. Your data stays inside your virtual boundary.

🕵️

Zero-Knowledge Architecture

AINNA cannot see your data. We manage the infrastructure, not your content. The VPN tunnel and VPS encryption ensure your data is opaque to everyone except your authorized agents.

📋

PDPA Compliance Built-In

Data sovereignty is the foundation of PDPA compliance. By keeping personal data within Malaysia and under your control, NeuralOps satisfies Section 129 (data transfer restrictions) automatically.

🛡️

Audit Trail Without Data Exposure

All inference requests are logged for compliance but only metadata (timestamp, token count, model used). The actual prompt content is never logged. You get auditability without exposure.

NeuralOps vs. The Alternatives

See how NeuralOps compares to OpenAI, Google Gemini, and self-hosted solutions across security, sovereignty, and control dimensions.

Swipe left/right to compare all options →
Feature NeuralOps OpenAI API Google Gemini Self-Hosted
VPN-Only Access ✅ Yes ❌ No ❌ No ⚠️ DIY
No Public API Endpoint ✅ Yes ❌ Public ❌ Public ✅ Yes
Whitelisted VPS Only ✅ Yes ❌ API key only ❌ API key only ⚠️ Manual
WireGuard Encryption ✅ Yes ⚠️ HTTPS only ⚠️ HTTPS only ⚠️ Your setup
Data Stays in Malaysia ✅ Yes ❌ US servers ❌ US/SG ✅ Your choice
PDPA Compliance Ready ✅ Built-in ❌ GDPR only ❌ GDPR only ⚠️ DIY
No Training on Your Data ✅ Guaranteed ⚠️ Opt-out ❌ May train ✅ Yes
Audit Trail ✅ Full ⚠️ Limited ⚠️ Limited ⚠️ DIY
Zero Data Retention ✅ Yes ❌ 30 days ❌ Varies ✅ Yes
Centralized Management ✅ AINNA manages ❌ OpenAI controls ❌ Google controls ❌ You manage

Defense in Depth Multi-Layer Security Architecture

A layered security chart showing how the public internet is kept outside while only designated VPS nodes can reach the vLLM core through VPN.

Governance & PDPA
Access Control
WireGuard VPN
Zero Open Ports
VPS Isolation
Snapshot Recovery
🧠
Central vLLM Inference Core
🌐 Public Internet
Blocked
🖥️ Designated VPS
Whitelisted
🤖 Agents
Controlled execution
Public traffic stops outside the security rings
Only whitelisted VPS enters through WireGuard VPN
vLLM remains an inference-only protected core

Frequently Asked Questions

Straight answers on models, VPN security, onboarding, SLA, and how NeuralOps compares to self-hosting.

🤖 Models 🔐 Security 🚀 Onboarding 📋 SLA & Support 💰 Cost
01 Models Which LLM models are available on the central vLLM server?

We run 7 open-weight models optimized for different use cases: Llama 3.1 (8B/70B), Qwen 2.5 (7B/32B/72B), Nemotron 3 Ultra, and Mistral Nemo 12B. Model availability varies by plan Starter gets 2 models, Business gets 5, Enterprise gets all 7. All models run locally on our GPU cluster; no external API calls.

02 Security What does "VPN-only access" actually mean for my team?

Your designated VPS receives a WireGuard config with a private IP (10.x.x.x). Only traffic originating from that VPS through the encrypted tunnel reaches the vLLM server. There is no public IP, no public DNS, no open ports on the LLM server. Your developers SSH into the VPS, run agents/apps there, and those apps call the private vLLM endpoint. The public internet cannot reach the inference layer at all.

03 Security Is my data used to train or improve the models?

Absolutely not. Zero data retention on the vLLM server prompts are processed in-memory and discarded immediately. No logging of prompt content (only metadata: timestamp, token count, model). Your VPS is isolated at hypervisor level; AINNA staff cannot access your files or memory. This is contractual and architectural.

04 Onboarding How long does onboarding take and what's included?

Starter: 3–4 days. Business/Enterprise: 5–7 days. Includes: VPS provisioning (Malaysia DC), WireGuard tunnel setup, vLLM model allocation, DNS + SSL for your app subdomain, SSH keys, monitoring agent install, and a 1-hour handover call. Enterprise adds: dedicated account engineer, custom SLA review, and Hermes gateway integration if needed.

05 Models What happens when I exceed my daily token limit?

Requests beyond your daily quota return a 429 response with a retry-after header. No overage charges the limit resets at 00:00 UTC. Enterprise plans support custom rate limits. You can monitor usage via the VPS dashboard or request a limit increase through your account engineer.

06 SLA What SLA do you offer and what's the uptime guarantee?

Enterprise: 99.9% uptime SLA with financial credits (pro-rata refund for downtime > 0.1%). Business: best-effort with priority support. Starter: community-tier monitoring. All tiers include: 24/7 infrastructure monitoring, automated failover for VPS layer, and snapshot-based recovery (RPO < 1 hour, RTO < 30 min).

07 Models Can I bring my own model or fine-tune on the central vLLM?

Enterprise plans support custom model deployment (GGUF / Safetensors) on dedicated GPU partitions requires security review and 2-week lead time. Fine-tuning is not offered on the shared vLLM; we recommend running fine-tuning jobs on your VPS (we provide GPU-enabled VPS add-ons) and deploying the adapter/merged model to your dedicated partition.

08 Cost How does NeuralOps compare to self-hosting vLLM on my own GPU server?

Self-hosting gives you full control but requires: GPU procurement (H200 lead times), 24/7 ops expertise, VPN + hardening, model optimization, monitoring, and compliance auditing. NeuralOps offloads all infrastructure ops you get a hardened, monitored, PDPA-compliant inference layer in days, not months. Cost comparison: RM3,500/mo (Enterprise) vs ~RM25K+/mo for equivalent self-hosted stack (GPU lease + colocation + engineering). Note: H200 GPU server lease alone (no colo/engineering) ranges RM20K–RM25K/mo the RM25K+ figure reflects the full managed stack.

09 Support What support channels are available for each tier?

Starter: Email/ticket (48h response). Business: Slack/Teams + ticket (4h business hours). Enterprise: Dedicated Slack channel + phone + ticket (1h response) + quarterly architecture review. All tiers include access to runbooks, API docs, and the AINNA status page.

10 Location Where are the servers physically located?

All infrastructure central vLLM GPU cluster and customer VPS instances is hosted in Tier III data centers in Cyberjaya and Kuala Lumpur, Malaysia. Data never leaves Malaysian jurisdiction. This satisfies PDPA Section 129 (cross-border transfer restrictions) by default.

Start Your Pilot Project

Start with one workflow. Build one system. Scale the intelligence layer from there.

🔒 Enterprise-grade AI infrastructure with VPN-only security. Private API only. Designed for controlled data exposure.

Research evidence: Research Hub · NeuralOps Architecture · Token Efficiency Benchmark

Get Started

Build AI around your operations.

Start with one workflow, department, or operational problem and scale through NeuralOps.

AINNA
CLICK ME

Site Sections

No section data available yet.

Sites with documented sections will appear here.

AINNA NeuralOps System