AI Strategy & Consulting
AI readiness assessment, use-case identification, enterprise roadmap, architecture planning, model selection, and smart routing strategy.
Discuss Deployment →NeuralOps Integrated Enterprise AI Architecture
NeuralOps connects private AI infrastructure, models, agents, data intelligence, automation, and detached systems behind a locked, governed network so intelligence runs inside your business, not around it.
Direct Answer
AINNA AI provides enterprise AI infrastructure, private AI, local LLM deployment, agentic AI, automation, data intelligence and governed NeuralOps patterns for Malaysian organisations that need practical delivery, not generic demos.
AI Services
A complete service stack strategy, private infrastructure, local models, agents, automation, data intelligence, and governed deployment.
AI readiness assessment, use-case identification, enterprise roadmap, architecture planning, model selection, and smart routing strategy.
Discuss Deployment →Private vLLM, VPS infrastructure, VPN-secured access, controlled inference, private AI gateway, and model hosting with infrastructure segmentation.
View Architecture →Local model deployment, evaluation, quantized models, multi-model architecture, model routing and segmentation, inference optimization, and lifecycle management.
Explore LLM Hub →Enterprise, operations, finance, customer-support, research, monitoring, and executive-intelligence agents plus governed multi-agent systems.
Explore Agent Hub →PLAN → DESIGN → BUILD → TEST → VALIDATE → HUMAN APPROVAL → DEPLOY → MONITOR, with audit trails and controlled execution.
Learn More →Workflow, business-process, and data automation; API/system integration; scheduled and event-driven operations; human-approval workflows; monitoring.
Learn More →AI builds → validate → detach → the system runs. Deterministic systems can continue operating without continuous LLM inference where appropriate.
Learn More →Data extraction, document parsing, classification, normalization, validation, reconciliation, and structured business data pipelines.
Learn More →Enterprise knowledge base, secure RAG, internal and semantic search, knowledge agents, and document retrieval with governed answers.
Learn More →AI-enabled web systems, business applications, internal dashboards, API and database systems, automation portals, and custom enterprise tools.
Start Pilot →ERP, CRM, ecommerce, database, API, and legacy-system integration plus agent-to-system integration and internal systems.
View Architecture →Human approval, RBAC, logging, audit trail, usage policy, model access control, data boundaries, monitoring, and operational guardrails.
Learn More →Smart routing, model segmentation, small-model routing, deterministic processing, detached systems, and inference optimization.
Explore LLM Strategies →Private network architecture, local AI, controlled access, VPN-based inference, VPS segmentation, logging, and policy enforcement.
View Sovereignty →Infrastructure, agent, workflow, model, and performance monitoring; incident detection; usage reporting; and system maintenance.
Start Pilot →NeuralOps Service Stack
Each layer builds on the one below it from private infrastructure to business outcome.
Industries
NeuralOps adapts to how each industry actually works with governed, private, and detached AI that fits the operational context.
Inventory intelligence, marketplace operations, sales analytics, pricing, order automation, and customer-service agents.
Bank statement processing, financial data extraction, reconciliation, reporting, accounting automation, and cash flow intelligence.
Production monitoring, predictive maintenance, quality control, machine data, inventory planning, and industrial AI agents.
City operations intelligence, infrastructure and traffic monitoring, utility management, and public service automation.
Institutional knowledge, secure internal AI, document intelligence, workflow automation, and audit & governance.
Document intelligence, administrative automation, knowledge retrieval, research assistance, and data governance.
Document processing, contract intelligence, knowledge retrieval, classification, and controlled human review.
Farm monitoring, crop and environmental data, resource optimization, sensor integration, and yield analytics.
Energy monitoring, grid and asset intelligence, consumption analytics, maintenance, and forecasting.
Shipment monitoring, route intelligence, warehouse automation, inventory visibility, and exception detection.
Building operations, maintenance automation, energy monitoring, asset management, and tenant service agents.
Branch monitoring, franchise compliance, sales intelligence, inventory monitoring, and commission automation.
Network operations, infrastructure monitoring, incident workflow, customer operations, and support automation.
Institutional knowledge, administrative automation, research assistance, and internal knowledge retrieval.
Project monitoring, technical knowledge, document processing, asset tracking, and procurement workflows.
Industry × AI Service Matrix
Select an industry to see the recommended NeuralOps services and the direct industry solution page.
| Service | Retail | Finance | Manufacturing | Government | Smart City |
|---|---|---|---|---|---|
| Private AI | ✓ | ✓ | ✓ | ✓ | ✓ |
| AI Agents | ✓ | ✓ | ✓ | ✓ | ✓ |
| Automation | ✓ | ✓ | ✓ | ✓ | ✓ |
| Detached Systems | ✓ | ✓ | ✓ | ✓ | ✓ |
| Data Intelligence | ✓ | ✓ | ✓ | ✓ | ✓ |
| Governance | ✓ | ✓ | ✓ | ✓ | ✓ |
The user never connects to the vLLM server directly. Requests go to the designated VPS first, then the VPS calls the central vLLM through a private WireGuard VPN tunnel.
Your business app, agents, dashboards, and automation run on an isolated VPS. This VPS becomes the only approved execution node for your AI workflow.
AINNA creates a WireGuard tunnel between your VPS and the central vLLM server. The VPS IP and tunnel credentials are whitelisted. Public traffic is rejected by design.
Agents send prompts through VPN, vLLM returns model output, and operations continue on the VPS. Public API keys are not required, and the model backend is not exposed in the default deployment model.
Complete AI infrastructure stack from local models to autonomous agents.
Private AI infrastructure running on-premise
Enterprise-grade security and data sovereignty
Autonomous agents that execute complex workflows
Design and deploy custom AI agents
Private AI infrastructure running on-premise
Enterprise-grade security and data sovereignty
Autonomous agents that execute complex workflows
Design and deploy custom AI agents
Enterprise-grade AI infrastructure with VPN-only security. Private deployment only. Designed for controlled data exposure.
For small businesses and startups
For growing companies and e-commerce
For enterprises and high-security needs
Not a private LLM per customer. One centralized vLLM server. Authorized VPS instances connect via VPN. Public-facing access is disabled in the default deployment. Hub-and-spoke architecture.
No per-client model isolation. Instead, a single optimized vLLM server shared via secure VPN tunnels more efficient and equally secure.
Single server running 7 LLM models. Fully utilized GPU resources. Centralized monitoring, updates, and security patches.
Only designated VPS instances with whitelisted VPN credentials can connect. Public API keys are not required. Public endpoints are not exposed in the default deployment model.
Inference ports remain private in the default deployment. Administrative access is restricted, the model backend is not publicly exposed, and direct GPU access from outside is blocked.
Central vLLM server connects to isolated VPS spokes via VPN. If one VPS fails, others remain unaffected. LLM server stays protected.
Every inference request logged. Centralized monitoring across all VPS connections. Complete auditability for compliance.
No per-client model isolation. Instead, a single optimized vLLM server shared via secure VPN tunnels more efficient and equally secure.
Single server running 7 LLM models. Fully utilized GPU resources. Centralized monitoring, updates, and security patches.
Only designated VPS instances with whitelisted VPN credentials can connect. Access is handled privately, and public endpoints are not exposed in the default deployment model.
Inference ports remain private in the default deployment. Administrative access is restricted, the model backend is not publicly exposed, and direct GPU access from outside is blocked.
Central vLLM server connects to isolated VPS spokes via VPN. If one VPS fails, others remain unaffected. LLM server stays protected.
Every inference request logged. Centralized monitoring across all VPS connections. Complete auditability for compliance.
Your data never leaves your controlled environment. All inference happens behind VPN. Third-party API calls are avoided where possible, and data exposure to public models is designed to be controlled. Security policy (kebijakan keselamatan) is enforced at the network perimeter.
Business data stays here. Agents process, agents execute. No data leaves your isolated zone.
Only prompt text travels through. Fully encrypted. No stored logs of your data.
Receives prompt → returns output. No data stored. No data logged. No data shared.
Supported bank-statement formats can be converted into structured financial records through dedicated parsers and validation. Sales, inventory, listings, advertising, logistics, and customer behaviour require source-specific connectors or custom parser integrations before downstream dashboards and models are built.
Extract insights from operational data
Retrieval-augmented generation with your data
High-performance semantic search and retrieval
AI-built systems that run independently
Extract insights from operational data
Retrieval-augmented generation with your data
High-performance semantic search and retrieval
AI-built systems that run independently
AINNA automates repetitive business processes across ecommerce, inventory, reporting, listing management, operations, and internal workflows. The goal is not just automation, but intelligent orchestration between data, rules, people, and systems.
End-to-end online store management
Real-time stock tracking and alerts
Automated revenue analytics and insights
ML-powered forecasting and trends
End-to-end online store management
Real-time stock tracking and alerts
Automated revenue analytics and insights
ML-powered forecasting and trends
Instead of relying only on chatbot interfaces, AINNA uses AI as a system builder designing, generating, repairing, and optimizing business applications that continue to operate independently after deployment.
Every system is designed for specific business outcomes, not generic conversations.
Your business logic, data, and insights remain under your control no external dependencies.
Systems continue running after deployment, requiring minimal maintenance and zero token costs for execution (infrastructure still applies separately).
Every system is designed for specific business outcomes, not generic conversations.
Your business logic, data, and insights remain under your control no external dependencies.
Systems continue running after deployment, requiring minimal maintenance and zero token costs for execution (infrastructure still applies separately).
AINNA develops AI Agents that assist in system planning, code generation, workflow design, data mapping, error detection, documentation, optimization, and system repair. These agents are not uncontrolled bots. They operate within clear business rules, human approval layers, and defined system boundaries.
Agents analyze business requirements and architect optimal system structures.
AI-assisted development produces clean, production-ready code.
Agents identify bugs, edge cases, and performance issues then fix them.
Every agent action follows business rules with human approval and audit trails.
Agents analyze business requirements and architect optimal system structures.
AI-assisted development produces clean, production-ready code.
Agents identify bugs, edge cases, and performance issues then fix them.
Every agent action follows business rules with human approval and audit trails.
Start at the secured platform hub, route through the multi-model orchestra, then deploy on-prem when data sovereignty is non-negotiable.
VPN-secured central vLLM
Agents and VPS nodes connect through whitelisted WireGuard tunnels no public inference endpoints.
You are here 027-model vLLM system
Route workloads across specialized models with centralized GPU efficiency and zero per-token cloud billing.
03Local models · zero data leakage
Run models inside your building when PDPA, sector rules, or air-gapped policy require full data residency.
The vLLM server is the private inference layer. It runs centrally for performance and model efficiency, but it is not visible to the public internet. Only designated VPS nodes can call it through encrypted VPN.
🔒 Private access only · Private inference ports · Restricted admin access · No direct GPU access · Whitelisted VPS only.
The vLLM server is reachable only from approved VPS nodes through private WireGuard tunnels.
The model backend is not directly reachable from the public internet.
Generic Agent AI, OpenCode, Detached Systems, and Hermes workflows run on VPS, not on the GPU server.
Central vLLM can serve multiple models while keeping access controlled through VPN-only routing.
The vLLM server is reachable only from approved VPS nodes through private WireGuard tunnels.
The model backend is not directly reachable from the public internet.
Generic Agent AI, OpenCode, Detached Systems, and Hermes workflows run on VPS, not on the GPU server.
Central vLLM can serve multiple models while keeping access controlled through VPN-only routing.
AI designs, develops, and validates the workflow once then the detached system runs on PHP, rules, databases, and automation without repeated inference.
Explore Detached Systems →AINNA's model is control-first. Human-in-the-loop approval, audit trail, system logging, access control, role-based permissions, monitoring, observability, and business rule enforcement are included to keep AI adoption safe, explainable, and manageable.
Human-in-the-loop approval, audit trail, system logging, access control, and business rule enforcement.
Every automated decision can be traced and explained with clear escalation paths.
Approval workflows for critical decisions with configurable risk levels.
Complete logging, observability dashboards, and compliance reviews.
Your data never leaves your controlled VPS environment. Security policy enforced at network perimeter. No third-party exposure.
Every automated decision can be traced and explained with clear escalation paths.
Approval workflows for critical decisions with configurable risk levels.
Complete logging, observability dashboards, and compliance reviews.
Citation-style summary based on verified operating facts from the live AINNA commerce environment.
AINNA's AI architecture evolved inside a live multi-store commerce environment with more than 80,000 active SKUs, 9,000 monthly orders, 30 official stores and RM15M+ in lifetime sales. These figures describe the operating environment, not AI performance.
Commerce operations required better control over catalogue breadth, order flow, store coordination, reporting and repeatable workflows across channels.
The environment involved multiple stores, frequent product changes, operational reporting pressure and the need to keep deterministic business rules visible.
NeuralOps routing, detached systems, controlled validation and private AI components were used to separate repeatable work from model-heavy work.
The AI layer handled language-heavy tasks, orchestration support and classification where flexible reasoning helped. It was not used as the only decision point.
Detached systems handled routing, validation, inventory logic, order checks and other repeatable business rules before action was taken.
The public evidence supports a practical AI operating model that fits real commerce pressure, not a purely experimental demo.
This is an operational case study, not a controlled academic experiment. The figures should not be read as universal AI benchmark claims.
Suggested citation: AINNA. "AINNA Operational Case Study." AINNA Research, 2026. See also Research Hub and NeuralOps Architecture.
Pitch 1 represents AINNA's expansion model combining business experience, AI infrastructure, detached system development, and data-driven execution into a scalable framework for SME transformation.
Single use case proof of concept with measurable ROI.
Expand to multiple teams and integrate with core systems.
Organization-wide AI ecosystem with full automation.
Scalable framework for SME transformation.
Single use case proof of concept with measurable ROI.
Expand to multiple teams and integrate with core systems.
Organization-wide AI ecosystem with full automation.
Scalable framework for SME transformation.
Secure AI infrastructure tailored to your industry's compliance and data sensitivity requirements.
Product description generation, customer service agents, inventory forecasting, and automated listing management across multiple stores.
Document analysis, contract review, case law research confidential data stays behind your VPN in the default deployment. No third-party exposure.
Patient data processing, clinical report generation, medical record analysis designed to support PDPA-aligned deployment controls, with data processing kept within the VPS in the default model.
Report generation, compliance monitoring, fraud detection, risk analysis secure, auditable, with complete inference logging.
Citizen services automation, document processing, policy analysis sovereign AI infrastructure with Malaysia-based data control.
IoT sensor monitoring, predictive maintenance, quality control automation, supply chain optimization with real-time agent execution.
NeuralOps protects sovereignty by separating user access, agent execution, and LLM inference. Users reach the VPS/app layer. Only the designated VPS reaches vLLM through VPN.
The user interacts with your app, dashboard, API, or agent running on VPS. The public side terminates here. The vLLM server is never exposed as a public destination.
The VPS uses WireGuard credentials and an approved IP route. If traffic does not originate from an authorized VPS tunnel, the vLLM layer rejects it.
The LLM server receives a controlled prompt over VPN and returns output. It is not used as a storage layer, web app layer, or public integration point.
All servers and VPS instances are hosted within Malaysia's borders. Your data is subject to Malaysian law (PDPA 2010), not foreign jurisdictions like US Cloud Act or GDPR.
Unlike solutions that proxy through OpenAI/Google/xAI APIs, NeuralOps makes zero external API calls. Your prompt never reaches a foreign server. Inference is 100% local to our infrastructure.
Each customer's VPS is isolated at the cloud hypervisor level. No other tenant can access your files, processes, or memory. Your data stays inside your virtual boundary.
AINNA cannot see your data. We manage the infrastructure, not your content. The VPN tunnel and VPS encryption ensure your data is opaque to everyone except your authorized agents.
Data sovereignty is the foundation of PDPA compliance. By keeping personal data within Malaysia and under your control, NeuralOps satisfies Section 129 (data transfer restrictions) automatically.
All inference requests are logged for compliance but only metadata (timestamp, token count, model used). The actual prompt content is never logged. You get auditability without exposure.
See how NeuralOps compares to OpenAI, Google Gemini, and self-hosted solutions across security, sovereignty, and control dimensions.
| Feature | NeuralOps | OpenAI API | Google Gemini | Self-Hosted |
|---|---|---|---|---|
| VPN-Only Access | ✅ Yes | ❌ No | ❌ No | ⚠️ DIY |
| No Public API Endpoint | ✅ Yes | ❌ Public | ❌ Public | ✅ Yes |
| Whitelisted VPS Only | ✅ Yes | ❌ API key only | ❌ API key only | ⚠️ Manual |
| WireGuard Encryption | ✅ Yes | ⚠️ HTTPS only | ⚠️ HTTPS only | ⚠️ Your setup |
| Data Stays in Malaysia | ✅ Yes | ❌ US servers | ❌ US/SG | ✅ Your choice |
| PDPA Compliance Ready | ✅ Built-in | ❌ GDPR only | ❌ GDPR only | ⚠️ DIY |
| No Training on Your Data | ✅ Guaranteed | ⚠️ Opt-out | ❌ May train | ✅ Yes |
| Audit Trail | ✅ Full | ⚠️ Limited | ⚠️ Limited | ⚠️ DIY |
| Zero Data Retention | ✅ Yes | ❌ 30 days | ❌ Varies | ✅ Yes |
| Centralized Management | ✅ AINNA manages | ❌ OpenAI controls | ❌ Google controls | ❌ You manage |
A layered security chart showing how the public internet is kept outside while only designated VPS nodes can reach the vLLM core through VPN.
Straight answers on models, VPN security, onboarding, SLA, and how NeuralOps compares to self-hosting.
We run 7 open-weight models optimized for different use cases: Llama 3.1 (8B/70B), Qwen 2.5 (7B/32B/72B), Nemotron 3 Ultra, and Mistral Nemo 12B. Model availability varies by plan Starter gets 2 models, Business gets 5, Enterprise gets all 7. All models run locally on our GPU cluster; no external API calls.
Your designated VPS receives a WireGuard config with a private IP (10.x.x.x). Only traffic originating from that VPS through the encrypted tunnel reaches the vLLM server. There is no public IP, no public DNS, no open ports on the LLM server. Your developers SSH into the VPS, run agents/apps there, and those apps call the private vLLM endpoint. The public internet cannot reach the inference layer at all.
Absolutely not. Zero data retention on the vLLM server prompts are processed in-memory and discarded immediately. No logging of prompt content (only metadata: timestamp, token count, model). Your VPS is isolated at hypervisor level; AINNA staff cannot access your files or memory. This is contractual and architectural.
Starter: 3–4 days. Business/Enterprise: 5–7 days. Includes: VPS provisioning (Malaysia DC), WireGuard tunnel setup, vLLM model allocation, DNS + SSL for your app subdomain, SSH keys, monitoring agent install, and a 1-hour handover call. Enterprise adds: dedicated account engineer, custom SLA review, and Hermes gateway integration if needed.
Requests beyond your daily quota return a 429 response with a retry-after header. No overage charges the limit resets at 00:00 UTC. Enterprise plans support custom rate limits. You can monitor usage via the VPS dashboard or request a limit increase through your account engineer.
Enterprise: 99.9% uptime SLA with financial credits (pro-rata refund for downtime > 0.1%). Business: best-effort with priority support. Starter: community-tier monitoring. All tiers include: 24/7 infrastructure monitoring, automated failover for VPS layer, and snapshot-based recovery (RPO < 1 hour, RTO < 30 min).
Enterprise plans support custom model deployment (GGUF / Safetensors) on dedicated GPU partitions requires security review and 2-week lead time. Fine-tuning is not offered on the shared vLLM; we recommend running fine-tuning jobs on your VPS (we provide GPU-enabled VPS add-ons) and deploying the adapter/merged model to your dedicated partition.
Self-hosting gives you full control but requires: GPU procurement (H200 lead times), 24/7 ops expertise, VPN + hardening, model optimization, monitoring, and compliance auditing. NeuralOps offloads all infrastructure ops you get a hardened, monitored, PDPA-compliant inference layer in days, not months. Cost comparison: RM3,500/mo (Enterprise) vs ~RM25K+/mo for equivalent self-hosted stack (GPU lease + colocation + engineering). Note: H200 GPU server lease alone (no colo/engineering) ranges RM20K–RM25K/mo the RM25K+ figure reflects the full managed stack.
Starter: Email/ticket (48h response). Business: Slack/Teams + ticket (4h business hours). Enterprise: Dedicated Slack channel + phone + ticket (1h response) + quarterly architecture review. All tiers include access to runbooks, API docs, and the AINNA status page.
All infrastructure central vLLM GPU cluster and customer VPS instances is hosted in Tier III data centers in Cyberjaya and Kuala Lumpur, Malaysia. Data never leaves Malaysian jurisdiction. This satisfies PDPA Section 129 (cross-border transfer restrictions) by default.
Start with one workflow. Build one system. Scale the intelligence layer from there.
🔒 Enterprise-grade AI infrastructure with VPN-only security. Private API only. Designed for controlled data exposure.
Research evidence: Research Hub · NeuralOps Architecture · Token Efficiency Benchmark
Get Started
Start with one workflow, department, or operational problem and scale through NeuralOps.