Photo: Tim Witzdam / Pexels
NEAR AI Cloud brings private inference to GitHub Copilot through a VSCode extension
A new VSCode extension lets GitHub Copilot run models from NEAR AI Cloud, where prompts and outputs are shielded inside hardware-secured enclaves
Developers who love GitHub Copilot but wince at sending their code to someone else’s servers now have another option. A new VSCode extension lets Copilot run models hosted on NEAR AI Cloud, a service built around private inference.
How the Copilot connection works
The integration runs through the “More Providers” VS Code extension, which plugs into GitHub Copilot Chat. That extension supports 258 OpenAI-compatible options, and NEAR AI Cloud is one of them.
NEAR AI Cloud slots in neatly because it exposes an OpenAI-compatible API. Tools built to talk to OpenAI-style endpoints can point at NEAR AI Cloud without rewiring existing workflows.
A June 2026 update to VS Code’s Bring Your Own Key (BYOK) support made this easier still. That change removed the GitHub sign-in requirement for custom and OpenAI-compatible endpoints.
What makes the inference private
The privacy claim rests on Trusted Execution Environments, or TEEs. NEAR AI Cloud builds these on Intel TDX and NVIDIA confidential computing technologies.
On NEAR AI Cloud, that protection covers three things: the prompts users submit, the model weights, and the outputs the model generates. All of them are isolated from infrastructure providers, model owners, and NEAR AI itself.
Each request also comes with a hardware-signed cryptographic attestation confirming the job ran inside a genuine secure environment.
AI, tech, and the markets they move—in one daily briefing.
Daily. Free. Join 34,000+ readers across crypto, finance, and policy.
Users who want extra assurance can turn on client-side encryption. With that option, data stays inaccessible to NEAR AI as well. Independent verification methods are also available.
Pricing, models, and a crypto-native payment route
NEAR AI Cloud hosts a growing lineup of open-weight models, including GLM 5.3 Flash. Input costs range from approximately $0.15 to $0.50 per million tokens, with output priced higher.
Users can stake NEAR tokens instead of paying with a credit card. The yield from that stake converts into inference credits, and the original stake stays intact when the user unstakes.
Background: a young platform moving quickly
NEAR AI Cloud launched in December 2025. In August 2026, the platform integrated the SayGm gateway. An OpenRouter integration followed in September 2026, widening the ways developers can reach the service.
What this means for developers and NEAR holders
For developers, the immediate value is choice. Teams with strict data policies can keep using Copilot Chat while pointing it at an inference provider built around hardware isolation and per-request attestations. The tradeoff is model selection: NEAR AI Cloud runs open-weight models, so teams will need to judge whether those models match the output quality they get elsewhere.
There is also the matter of trust shifting rather than disappearing. TEEs reduce reliance on the cloud operator, but they place confidence in the chip vendors’ security designs. The attestations and client-side encryption options exist precisely to give users ways to verify rather than simply trust.
For NEAR token holders, the staking model ties inference credits to staking yield, giving the token a concrete job inside the platform.