NVIDIA-Certified Associate Generative AI LLMs Practice Questions
Prepare for NCA-GENL with more than an answer.
- Exam fee
- $125 USD
- Level
- Associate
- Valid for
- 2 years
Domains covered on the exam 5
- Core Machine Learning and AI Knowledge30%
- Software Development24%
- Experimentation22%
- Data Analysis and Visualization14%
- Trustworthy AI10%
- 1
A startup is building a customer support chatbot. They have limited GPU resources and cannot afford to fully fine-tune a large language model. They need a method to adapt a pre-trained model to their company's specific support documents and tone. Which of the following is the MOST resource-efficient fine-tuning strategy?
Show answer details
Correct answer: B
Parameter-Efficient Fine-Tuning (PEFT) methods, such as LoRA (Low-Rank Adaptation), are designed for this exact scenario. They freeze the vast majority of the pre-trained model's parameters and only train a small number of additional parameters (adapters). This dramatically reduces memory and compute requirements, making it possible to fine-tune large models on consumer-grade or limited enterprise hardware.
- 2
An organization is deploying a generative AI application using the recently announced NVIDIA NIM (NVIDIA Inference Microservices). What is the primary advantage of using NIM for deployment?
Show answer details
Correct answer: C
NVIDIA NIM (NVIDIA Inference Microservices) provides pre-built, cloud-native microservices that simplify the deployment of generative AI models. It packages models and their dependencies into optimized containers, exposing a standard industry API. This abstracts away the complexity of the underlying inference stack (like TensorRT-LLM and Triton), allowing developers to deploy highly performant models much more quickly and easily across various environments.
- 3
Case Study:
Company Background:
Global-Retail Inc. is an e-commerce giant planning to deploy a new AI-powered product recommendation engine. The engine will use a generative LLM to create personalized, descriptive product suggestions for millions of users in real-time. The goal is to increase user engagement and sales by providing compelling, human-like recommendations.Current Situation:
The AI team has fine-tuned a powerful open-source LLM on their product catalog. However, initial tests reveal two major problems. First, the inference latency is highly variable and often exceeds the 100ms target required for a good user experience. Second, the cost of the GPU infrastructure required to serve the model at scale is projected to be prohibitively expensive. The current deployment uses a basic Python Flask server with the Hugging Face Transformers library.Requirements:
- Reduce average inference latency to below 100ms.
- Maximize GPU utilization to improve throughput and reduce the number of required servers.
- Ensure the deployment is robust and can scale automatically to handle traffic spikes during sales events.
- Standardize the deployment process for future AI models.
Proposed Solution:
The lead architect proposes a new deployment architecture. The plan involves taking the fine-tuned model, optimizing it for NVIDIA GPUs, and serving it through a dedicated inference server that can handle concurrent requests efficiently.Which combination of NVIDIA technologies best addresses all of Global-Retail's requirements?
graph TD subgraph Current_State ["Current State (High Latency & Cost)"] UserRequest --> FlaskApp[Flask App] FlaskApp --> HF_Model[Hugging Face Model on GPU] end subgraph Proposed_Solution ["Proposed NVIDIA Solution"] LB[Load Balancer] --> Triton_Cluster[Triton Inference Server Cluster] Triton_Cluster --> TRT_LLM_Model[Optimized Model w/ TensorRT-LLM] TRT_LLM_Model -- In-flight Batching --> High_Throughput_GPU[High Throughput on GPU] end Current_State -->|Migration| Proposed_SolutionShow answer details
Correct answer: B
This solution directly addresses all the requirements. TensorRT-LLM optimizes the model for NVIDIA hardware, significantly reducing latency (Requirement 1). NVIDIA Triton Inference Server is a production-grade server that provides features like in-flight batching to maximize GPU utilization and throughput (Requirement 2). Triton is designed for scalable, robust deployments (Requirement 3) and provides a standardized method for serving models (Requirement 4).
- 4
During the fine-tuning of an LLM, a data scientist observes that the training loss continues to decrease, but the validation loss starts to increase after a certain number of epochs. What is this phenomenon called?
Show answer details
Correct answer: B
Overfitting occurs when a model learns the training data too well, including its noise and specific patterns, to the point where it performs poorly on new, unseen data. The classic sign of overfitting is when the training loss continues to decrease while the validation loss begins to rise, indicating a drop in generalization performance.
- 5
What are the three core components that define a self-attention calculation in a Transformer model? (Select THREE)
Show answer details
Correct answer: A, B, C
The Query vector represents the current word/token being processed and is used to score against all other keys.
The Key vectors are associated with all words in the sequence and are matched against the query to determine relevance.
The Value vectors contain the actual information/representation of each word. The attention scores (from Q-K dot products) are used to create a weighted sum of the value vectors.
- 6
True or False: The BERT architecture is an example of an autoregressive, decoder-only model.
Show answer details
Correct answer: B
The statement is false. BERT (Bidirectional Encoder Representations from Transformers) is an encoder-only model. It is designed to understand context by looking at both the left and the right side of a token simultaneously, making it non-autoregressive. Autoregressive, decoder-only models, like the GPT family, generate text one token at a time based only on the preceding tokens.
- 7
A developer needs to create a prompt that provides multiple examples of input-output pairs to guide the LLM's response format and style. This prompting technique is known as:
Show answer details
Correct answer: C
Few-shot prompting involves providing a few examples (shots) of the task within the prompt itself. This demonstrates the desired input-output pattern to the model, enabling it to perform the task more accurately for a new input, a process also known as in-context learning.
- 8
A data science team is preparing a large text dataset for fine-tuning a Llama 3 model. The dataset consists of 500GB of raw text files. The team needs to perform tokenization and data cleaning as quickly as possible. Which NVIDIA library is specifically designed for GPU-accelerated data manipulation and would be most suitable for this task?
Show answer details
Correct answer: C
NVIDIA RAPIDS cuDF is the correct choice. It provides a pandas-like API for data manipulation that runs on GPUs, making it ideal for accelerating data preprocessing tasks like cleaning and tokenization on large datasets. Triton is for inference serving, TensorRT-LLM is for optimizing inference, and NeMo is a framework for building and training models.
- 9
A developer is implementing a Retrieval-Augmented Generation (RAG) system to answer questions about internal company documents. They have already generated embeddings and stored them in a vector database. Which step in the RAG pipeline immediately follows the retrieval of relevant document chunks from the vector database?
Show answer details
Correct answer: B
In a standard RAG pipeline, after the system retrieves the most relevant document chunks (context) from the vector database based on the user's query, the next step is to combine this retrieved context with the original user prompt. This new, augmented prompt is then sent to the LLM to generate a contextually-aware answer. Generating the final answer happens after the prompt is augmented.
- 10
An MLOps engineer is deploying a large language model using NVIDIA Triton Inference Server. They observe that under high load, requests with long sequences are causing head-of-line blocking, increasing latency for all subsequent requests. Which Triton feature is specifically designed to mitigate this issue by processing requests out of order?
Show answer details
Correct answer: C
In-flight batching (also known as continuous batching) is the correct feature. Unlike dynamic batching, which waits to form a complete batch before processing, in-flight batching can add new requests to a batch that is already being processed. This allows shorter requests to be processed and return while longer ones are still running, effectively eliminating head-of-line blocking and improving overall throughput and latency.
