NEAR AI Cloud joins OpenRouter as provider, serving GLM 5.3 Flash with 1M context

Editorial illustration: A translucent blue cloud containing a row of cream-colored plates connects through a metal coupling to a dark cylindrical hub with three branching blue channels.

In brief

  • NEAR AI Cloud went live on OpenRouter under the provider slug near-ai.
  • Z.ai's GLM 5.3 Flash is served with a 1 million token context window.
  • Pricing is $0.15 per million input tokens and $0.50 per million output tokens.
  • Inference runs inside Intel TDX and NVIDIA confidential computing enclaves, per Crypto Briefing.
  • OpenRouter users need no NEAR AI account or NEAR tokens.

What GLM 5.3 Flash is

GLM 5.3 Flash was built by Z.ai. It launched publicly on August 26 after an anonymous preview period called "Ox Alpha," according to Crypto Briefing (the report didn't give a year).

The specs are big. Crypto Briefing describes it as a multimodal mixture-of-experts model with 320 billion total parameters, of which 18 billion are active at any given time. It takes text, images and video as inputs and produces text outputs. The open weights are released under the MIT license, per the same report. LeoDex News hasn't independently checked these figures, the context window or the pricing against a Z.ai model card or the OpenRouter listing.

The privacy pitch

This is where NEAR is trying to stand apart from other inference providers. NEAR AI Cloud runs inference inside hardware-based Trusted Execution Environments, specifically Intel TDX and NVIDIA's confidential computing solutions, the publication reported. Every request comes back with a cryptographic attestation report.

Crypto Briefing calls that report proof that the computation ran inside a secure enclave and that no one, NEAR included, could have accessed the data. That's the publication's characterization, not an independently verified finding. NEAR AI Cloud also keeps a zero data retention policy, so prompts and responses aren't stored or logged on NEAR's side, according to the report.

The report didn't cite an independent audit of the enclave setup or the retention policy.

Where the NEAR token fits

For developers already on OpenRouter, the friction is low. They can select the near-ai provider without creating a separate NEAR AI account, and they don't need to hold any NEAR tokens to use it.

The token still has a role, but it isn't a separate route in. NEAR holders who stake can convert their staking yields into compute credits for NEAR AI Cloud services, Crypto Briefing reported. That ties staking yield to compute spend for holders, while OpenRouter users can ignore the token entirely.

Frequently asked questions

How does NEAR AI Cloud protect prompts sent through OpenRouter?

According to Crypto Briefing, NEAR AI Cloud runs inference inside hardware-based Trusted Execution Environments (Intel TDX and NVIDIA's confidential computing solutions). Every request comes with a cryptographic attestation report. The platform also keeps a zero data retention policy, so prompts and responses aren't stored or logged on NEAR's side.

Do I need NEAR tokens to use GLM 5.3 Flash on OpenRouter?

No. According to Crypto Briefing, OpenRouter users can select the near-ai provider without creating a separate NEAR AI account and without holding any NEAR tokens. NEAR holders who stake can separately convert staking yields into compute credits for NEAR AI Cloud services.

How much does GLM 5.3 Flash cost on NEAR AI Cloud?

Crypto Briefing reported that GLM 5.3 Flash costs $0.15 per million input tokens and $0.50 per million output tokens on NEAR AI Cloud. The model, served with a 1 million token context window, was developed by Z.ai.