Sign in to the admin console
Unified inference infrastructure for production-grade LLM workloads. Sub-100ms latency, enterprise security, usage-based pricing — all behind a single OpenAI-compatible endpoint.
Live status of the backing GPU endpoint
Endpoint
—Name
—
GPU Type
—
GPUs
—
Workers Idle
—
Workers Running
—
Workers Pending
—
Worker Count
—
Image
—
Memory
—
Disk
—
Add this entry to your VS Code chatLanguageModels.json file to use NexusAI with GitHub Copilot:
[ { "name": "NexusAI", "vendor": "ollama-models", "url": "https://agent.ash-api.online" } ]
~/Library/Application Support/Code/User/chatLanguageModels.json
Live Monitoring
Monitor every API request, response, tokens consumed, and latency metrics in real time.
| Time | User | Endpoint | Model | Prompt | Tokens | Latency | View |
|---|