Nemotron 3.5 Lightning
Access Nemotron 3.5 Lightning through one API key. NVIDIA Open High-Throughput MoE.
NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA with 3B active parameters out of 30B total, suited for high-throughput agentic workloads and specialised tasks. Lightning is for throughput-bound agent pipelines and specialised tasks where an open model with a long context at very low cost is the goal.
No credit card required. Takes 30 seconds.
2,400+
Developers
1.2M+
API calls served
100+
Endpoints
$0.01
Per call
Yep, that's it.
Try it live
Send a message and see Nemotron 3.5 Lightning respond in real time.
Maximum tokens in the response.
Real-time tokens
Context Window
1M tokens
Max Output
131K tokens
Input Price
$0.12 / 1M tokens
Output Price
$0.30 / 1M tokens
Strengths
NVIDIA's open 30B MoE with 3B active parameters.
Designed for agentic workloads that run at scale.
Handles up to 1,000,000 input tokens and returns up to 131,072 output tokens per call.
$0.12 per 1M input and $0.30 per 1M output tokens.
Quick start
Copy this snippet and start making calls with Nemotron 3.5 Lightning.
const res = await fetch('https://api.yepapi.com/v1/ai/chat', {
method: 'POST',
headers: {
'x-api-key': 'YOUR_API_KEY',
'Content-Type': 'application/json',
},
body: JSON.stringify({
"model": "nvidia/nemotron-3.5-lightning",
"messages": [
{
"role": "user",
"content": "Explain API gateways in 2 sentences."
}
],
"maxTokens": 256
}),
});
const { data } = await res.json();
console.log(data.message.content);Why use Nemotron 3.5 Lightning through YepAPI?
Nemotron 3.5 Lightning API — pricing, context window & access
NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA with 3B active parameters out of 30B total, suited for high-throughput agentic workloads and specialised tasks. It pairs a 1,000,000-token context window with up to 131,072 output tokens per response, at $0.12 per 1M input and $0.30 per 1M output tokens through YepAPI.
Through YepAPI you call Nemotron 3.5 Lightning on an OpenAI-compatible endpoint with a single key — the same key that reaches GPT-6, Claude, Gemini, Grok and the SEO, SERP and scraping APIs. Input: text.
What is Nemotron 3.5 Lightning?
NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA with 3B active parameters out of 30B total, suited for high-throughput agentic workloads and specialised tasks. It offers a 1,000,000-token context window and up to 131,072 output tokens at $0.12 / $0.30 per 1M tokens. It accepts text as input, handles up to 1,000,000 tokens of context and returns up to 131,072 output tokens per call.
Build with Nemotron 3.5 Lightning via YepAPI
YepAPI serves Nemotron 3.5 Lightning on the OpenAI-compatible /v1/ai/chat endpoint. Point your base URL at YepAPI, add your key, and set the model string to nemotron-3.5-lightning (or the full nvidia/nemotron-3.5-lightning). Function calling, structured outputs and reasoning settings pass straight through, and switching to any other model on the platform is a one-string change per request.
Nemotron 3.5 Lightning API pricing — $0.12 / 1M input, $0.30 / 1M output
Nemotron 3.5 Lightning costs $0.12 per 1M input tokens and $0.30 per 1M output tokens through YepAPI, pay per token with no minimums. Failed calls are never charged, and every request is itemised in your dashboard API Logs so you can see exactly what a 131,072-token response cost.
Nemotron 3.5 Lightning use cases
Lightning is for throughput-bound agent pipelines and specialised tasks where an open model with a long context at very low cost is the goal.
Try Nemotron 3.5 Lightning free
Every new YepAPI account includes $5 in free credit with no card required — enough to test Nemotron 3.5 Lightning on your own prompts before committing to paid usage. Create an account, copy your key, and call nemotron-3.5-lightning right away.
Start generating in 30 seconds
$5 free credit on signup. No credit card required. Pay per call.
What developers say
“Switched from SerpAPI and cut our SERP costs by 80%. Same data quality, way simpler billing.”
“One API key for AI models, SERP data, and web scraping. Saved us from managing 4 separate providers.”
“The $5 free credit let us prototype our entire rank tracking feature before committing. No other API does that.”
Frequently asked questions
NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA with 3B active parameters out of 30B total, suited for high-throughput agentic workloads and specialised tasks. Lightning is for throughput-bound agent pipelines and specialised tasks where an open model with a long context at very low cost is the goal.
Input tokens cost $0.12 per 1M tokens and output tokens cost $0.30 per 1M tokens through YepAPI. No monthly minimums — you only pay for what you use.
Sign up for a free API key, then send requests to the /v1/ai/chat endpoint.
Nemotron 3.5 Lightning supports a 1M token context window with up to 131K output tokens per request.
Ready to use Nemotron 3.5 Lightning?
$5 free credit on signup. No credit card required. Pay per call.
Explore more models
Nemotron 3 Super
NVIDIAAccess Nemotron 3 Super through one API key. NVIDIA's free model.
GPT-4o Mini
OpenAIAccess GPT-4o Mini through one API key. Fast, cheap, and OpenAI-compatible.
GPT-4o
OpenAIAccess GPT-4o through one API key. Flagship reasoning and multimodal capabilities.