

NEAR AI Cloud
#19 in LLM-Inferenz & ServingNear Ai · 7× · tolest 30. Juni 2026
0
Momentum
NEAR AI Cloud is an Inference platform by NEAR AI that runs LLMs (including GLM, Qwen, GPT-OSS, Llama, DeepSeek, Mixtral) within hardware-based Trusted Execution Environments (Intel TDX + NVIDIA Confidential Computing). The API is OpenAI-compatible, and each request receives cryptographic attestation as proof that code and data ran unaltered within the secured enclave. The platform is currently in beta, with usage-based pricing per Token; Enterprise plans offer reserved capacity and private models. Official launch was in early December 2025 alongside NEAR Private Chat.
Momentum-Verloop
19.05.17.08.
Features
| Throughput/Latency | Supports streaming, fine-tuning, and high-throughput performance with minimal latency (typically 5 to 10 percent overhead) |
| License | Cloud API repository licensed under PolyForm Strict License 1.0.0; enclave runtime, verifier and model images are open and reproducible |
| Platform | Cloud API with hardware-secured enclaves (Intel TDX + NVIDIA Confidential Computing); currently in beta |
| Price | Pay-per-token, live per-model pricing (e.g. GLM-4.6 FP8: $0.75/M input tokens, $2/M output tokens); Enterprise with reserved capacity |
| Protocol Compatibility | OpenAI-compatible API; existing OpenAI SDK code works by simply changing the base URL |
| Release Date | December 3, 2025 (official launch of NEAR AI Cloud and NEAR Private Chat) |
| Supported Models/Providers | Catalog of open models such as Llama, Qwen, DeepSeek, Mixtral, GLM, GPT-OSS (TEE-hosted); plus proxied third-party providers like OpenAI, Anthropic, Gemini |