Nemotron 3 Ultra Nvfp4: capabilities, context and price
Nemotron 3 Ultra Nvfp4 is a Fireworks route represented in Tavory for nemotron-3-ultra-550b-a55b-nvfp4 is a frontier-scale large language model (llm) trained by nvidia, designed to deliver strong agentic, reasoning, and conversational capabilities. it is optimized for the most demanding workloads, including complex multi-step agents, long-context analysis, and high-accuracy reasoning over code, math, and science. the model employs a hybrid latent mixture-of-experts (latentmoe) architecture, utilizing interleaved mamba-2 and moe layers, along with select attention layers. like the super model, the ultra model incorporates multi-token prediction (mtp) layers for faster text generation and improved quality, and it is trained using an nvfp4 pre-training recipe to maximize compute efficiency. the model has 55b active parameters and 550b parameters in total..
Tavory's live catalog lists Nemotron 3 Ultra Nvfp4 through Fireworks. The route is described as nemotron-3-ultra-550b-a55b-nvfp4 is a frontier-scale large language model (llm) trained by nvidia, designed to deliver strong agentic, reasoning, and conversational capabilities. it is optimized for the most demanding workloads, including complex multi-step agents, long-context analysis, and high-accuracy reasoning over code, math, and science. the model employs a hybrid latent mixture-of-experts (latentmoe) architecture, utilizing interleaved mamba-2 and moe layers, along with select attention layers. like the super model, the ultra model incorporates multi-token prediction (mtp) layers for faster text generation and improved quality, and it is trained using an nvfp4 pre-training recipe to maximize compute efficiency. the model has 55b active parameters and 550b parameters in total.. This page summarizes the current catalog facts so you can compare context, capabilities and provider price signals before opening a workspace request. Availability, plan access and final cost are checked again when you send.
Good fit for
text and chat
reasoning and analysis
Strengths and limitations
Catalog strengths
Nemotron-3-Ultra-550B-A55B-NVFP4 is a frontier-scale large language model (LLM) trained by NVIDIA, designed to deliver strong agentic, reasoning, and conversational capabilities. It is optimized for the most demanding workloads, including complex multi-step agents, long-context analysis, and high-accuracy reasoning over code, math, and science. The model employs a hybrid Latent Mixture-of-Experts (LatentMoE) architecture, utilizing interleaved Mamba-2 and MoE layers, along with select Attention layers. Like the Super model, the Ultra model incorporates Multi-Token Prediction (MTP) layers for faster text generation and improved quality, and it is trained using an NVFP4 pre-training recipe to maximize compute efficiency. The model has 55B active parameters and 550B parameters in total.
4/5 catalog quality signal
4/5 catalog speed signal
Keep in mind
Live availability and plan access can change.
Catalog signals do not guarantee a result for every prompt.
How to use Nemotron 3 Ultra Nvfp4 in Tavory
01
Open the workspace
Sign in, start a conversation and keep the task or project context together.
02
Choose Nemotron 3 Ultra Nvfp4
Select the model manually so Tavory preserves your choice for the request.
03
Review the result
Check the displayed model, answer, usage and settled cost before continuing.
Tavory's live catalog lists Nemotron 3 Ultra Nvfp4 through Fireworks. The route is described as nemotron-3-ultra-550b-a55b-nvfp4 is a frontier-scale large language model (llm) trained by nvidia, designed to deliver strong agentic, reasoning, and conversational capabilities. it is optimized for the most demanding workloads, including complex multi-step agents, long-context analysis, and high-accuracy reasoning over code, math, and science. the model employs a hybrid latent mixture-of-experts (latentmoe) architecture, utilizing interleaved mamba-2 and moe layers, along with select attention layers. like the super model, the ultra model incorporates multi-token prediction (mtp) layers for faster text generation and improved quality, and it is trained using an nvfp4 pre-training recipe to maximize compute efficiency. the model has 55b active parameters and 550b parameters in total.. This page summarizes the current catalog facts so you can compare context, capabilities and provider price signals before opening a workspace request. Availability, plan access and final cost are checked again when you send.
What can Nemotron 3 Ultra Nvfp4 be used for?
Nemotron 3 Ultra Nvfp4 is represented in Tavory for text and chat, reasoning and analysis. Suitability still depends on the prompt and required capabilities.
Can I use Nemotron 3 Ultra Nvfp4 in Tavory?
If Nemotron 3 Ultra Nvfp4 is eligible for your plan and a healthy route is available when you send, you can choose it in Tavory. Tavory is an independent workspace and is not the manufacturer of Nemotron 3 Ultra Nvfp4.
How much does Nemotron 3 Ultra Nvfp4 cost in Tavory?
$0.6 input / $2.4 output per 1M tokens is the current public catalog signal. Tavory checks the selected route and shows the applicable request cost; provider data can change.