Skip to content

38+ Free AI Models With NVIDIA NIM API

If your AI subscription is running out too fast, doesn’t have enough variety, or you just want to try something new, NVIDIA is letting you run models for free through their API.

We’re talking everything from small 4B translation models all the way up to the 552B DeepSeek V4.1 Flash and the 753B GLM 5.3. Models that would normally require a rack of GPUs to run at home are suddenly just an API key away.

What is NVIDIA NIM?

NIM (NVIDIA Inference Microservices) is NVIDIA’s way of packaging AI models so they’re easy to deploy and run. On their NVIDIA Build site, you can browse a large catalog of models, many of them optimized to run on NVIDIA hardware, and try them directly in the browser or call them through an API.

The API is OpenAI-compatible, which means most tools that already support OpenAI-style endpoints can talk to it with very little setup. That’s what makes it so easy to plug into coding assistants like OpenCode.

In the catalog you’ll see a few different labels:

  • Free Endpoint – NVIDIA hosts the model and you can call it through the API for free.
  • Partner Endpoint – the model is hosted by one of NVIDIA’s partners.
  • Downloadable – you can download the model and run it on your own hardware.

Some highlights

A few of the models that caught my eye:

DeepSeek V4.1 Flash – a 552B mixture-of-experts model with only 8B active parameters, native multimodal support, and a smaller KV cache that keeps costs down. A good all-rounder for chat and coding.

GLM 5.3 – a 753B text model from Z.ai with DeepSeek-style sparse attention, native FP8 weights, reasoning, and tool calling. Great if you want something that can plan and use tools.

Kimi K3 – a roughly 2.8T parameter multimodal model from Moonshot, built for long-horizon coding, agentic tool use, and image understanding. Both downloadable and available as a free endpoint.

Riva Translate 4B Instruct v2 – a small, fast translation model from NVIDIA that handles 37 languages and supports few-shot example prompts.

free ai api

Having this many options in one place is great for experimenting. You can test the same prompt across several models and see which one actually works best for your task, instead of being locked into one provider.

You can access with just a few steps.

  1. Install OpenCode or similar tool that supports custom providers.
  2. Create an account at NVIDIA Build
  3. If prompted, send a short email to NVIDIA support to verify your email. This may take a little while, so be patient.
  4. Log in to NVIDIA Build and click your profile in the upper right corner
  5. Click on “API” and generate an API key. Copy it somewhere safe, since you’ll need it in the next step.
  6. In OpenCode menu, click “providers” and search for NVIDIA
  7. Paste your API and connect
Opencode nvidia

Once connected, NVIDIA shows up under your connected providers and you can pick any of the available models from the model list. You now have access to over 100 models, some through the API and others through downloads. Use the filters on NVIDIA Build to see which ones have a free endpoint and which are available for download.

all nvidia free models

A few things to keep in mind

  • Free doesn’t mean unlimited – There are rate limits, so this works best for personal projects, testing, and experimenting rather than heavy production use. Check NVIDIA’s current terms for the details.
  • Your prompts go to NVIDIA’s servers – Avoid sending anything sensitive, like passwords, customer data, or private code you’re not allowed to share.
  • Keep your API key private – Don’t paste it into public repos or share screenshots where it’s visible.
  • The catalog changes – Models get added and removed over time, so if one disappears, there’s usually a newer one to try.

Sign up for our newsletter to get more tips and ideas directly in your inbox!

Published inAIEnglishTech