Adaptive Practice

Question 1 of 20 · Questions 1–10 are free

You operate a NeMo-based agent that performs RAG over a large vector store and then queries an LLM accelerated with TensorRT-LLM behind Triton. Costs are rising and throughput is capped. Which configuration change best improves throughput per dollar while keeping response quality stable?