GEN-AI-ENG Generative AI Engineer Associate Practice Questions
Prepare for GEN-AI-ENG with more than an answer.
- Exam fee
- $200 USD
- Level
- Associate
- Valid for
- 2 years
Domains covered on the exam 6
- Design Applications14%
- Data Preparation14%
- Application Development30%
- Assembling and Deploying Applications22%
- Governance8%
- Evaluation and Monitoring12%
- 1
A software company maintains its internal documentation in a series of Markdown files. They are building a RAG chatbot to help new developers find information. The quality of the chatbot's answers is poor, and investigation reveals that code blocks and tables within the Markdown are being chunked improperly, breaking their syntax and context. Which chunking strategy would best preserve the integrity of these structured elements?
Show answer details
Correct answer: B
Standard text splitters are unaware of the document's structure. For formats like Markdown, a specialized, language-aware splitter is necessary. A Markdown text splitter can identify structural elements like headers, lists, code blocks (e.g., ```), and tables. It can then use these elements as logical boundaries for splitting, ensuring that semantically coherent units like an entire code block or table are not broken apart, which dramatically improves retrieval quality.
- 2
A gaming company is creating an NPC (non-player character) whose dialogue is generated by an LLM. To make the NPC's responses more engaging and context-aware, the developers want to augment the prompt with the player's recent actions, current inventory, and location. This information is available through various game engine APIs. Which architectural pattern is best suited for this task?
Show answer details
Correct answer: B
This scenario requires dynamic, real-time data that cannot be stored in a static vector database. The agent/tool-use pattern is perfect for this. An agent can be equipped with tools that are essentially wrappers around the game engine APIs. Before generating dialogue, the agent uses these tools to fetch the current, dynamic context and then constructs a rich prompt for the LLM, leading to highly relevant and immersive responses.
- 3
A startup is building a GenAI application that is exposed to the public. They are concerned about prompt injection attacks where a malicious user could input instructions that hijack the model's purpose, such as revealing the system prompt or ignoring previous instructions. Which TWO techniques are effective defenses against this? (Choose two.)
Show answer details
Correct answer: A, C
- 4
A developer needs to register a trained scikit-learn model to the Unity Catalog using MLflow. The model object is
model, the conda environment isconda_env, and the desired registered name isprod_catalog.ml_models.churn_predictor. Which MLflow command correctly registers the model?Show answer details
Correct answer: A
This is the correct syntax.
mlflow.sklearn.log_modelis the function for scikit-learn models. Theregistered_model_nameparameter is used to specify the three-level namespace (catalog.schema.model) for registering the model in Unity Catalog. Providing the conda environment ensures reproducibility. - 5
What is the primary difference between the evaluation phase and the monitoring phase in the GenAI application lifecycle?
flowchart TD subgraph Development A[Data Prep] --> B[Model/Prompt Dev] B --> C{Evaluation} C -- Good --> D[Deployment] C -- Bad --> B end subgraph Production D --> E[Live Application] E --> F{Monitoring} F -- Drift/Issues --> A endShow answer details
Correct answer: A
This is the core distinction. Evaluation is a pre-deployment, experimental phase. It uses a fixed, curated dataset (an 'evaluation set') to compare different versions of a model, prompt, or RAG strategy to select the best one. Monitoring is a post-deployment, operational phase. It tracks the application's performance on live, unseen user data to detect problems like data drift, performance degradation, or increased toxicity, which then triggers a new development cycle.
- 6
A financial services firm is developing a RAG application to answer analyst questions about quarterly earnings reports. The reports are dense PDFs. During evaluation, the team notices that retrieval often fails to find specific numerical data mentioned deep within tables. The current chunking strategy is a simple recursive character split with a size of 1000 and an overlap of 200. What is the most effective approach to improve retrieval accuracy for tabular data?
Show answer details
Correct answer: B
Simple text-based chunking strategies are ineffective for structured data like tables. The best approach is to use a specialized library to extract the tables, preserve their structure by converting them to a text-friendly format like Markdown, and then embed them. This ensures the semantic relationship between rows and columns is maintained, leading to much higher retrieval accuracy for queries about specific data points within those tables.
- 7
A legal tech company is building a RAG system using a large corpus of internal legal documents. To comply with data privacy regulations, any document containing Personally Identifiable Information (PII) must be handled with strict access controls. The documents are stored in Delta Lake and managed by Unity Catalog. How should an engineer implement a governance strategy to prevent unauthorized access to sensitive documents during the retrieval process?
Show answer details
Correct answer: B
The most robust and scalable governance strategy is to enforce security at the data source. By using Unity Catalog's built-in features like row-level security and column masking on the Delta tables, access control is managed centrally and consistently. The RAG application, operating under a service principal with limited permissions, will inherit these restrictions, ensuring that it can only retrieve and process data the user is authorized to see, thus enforcing compliance.
- 8
An e-commerce company has deployed a customer support chatbot. The team wants to use an LLM-as-a-judge approach to evaluate the helpfulness of the chatbot's responses. They have a dataset of customer queries but lack a corresponding set of human-written, 'golden' answers. Which evaluation method is most suitable in this scenario?
Show answer details
Correct answer: A
When ground truth (golden answers) is not available, pairwise comparison is a highly effective evaluation method. It relies on relative judgment rather than absolute correctness. The judge LLM's task is simplified to determining which of two responses is better, which is a more reliable and consistent task for an LLM than assigning an absolute score without a reference. This approach allows for effective ranking and selection of better-performing models or prompts without needing a costly human-labeled dataset.
- 9
A developer is building a RAG application that sources information from public websites. To keep the information current, the data ingestion pipeline runs daily, scraping new articles. The developer notices that many articles contain large, irrelevant sections like advertisements, navigation menus, and user comments, which are degrading the quality of the retrieved context. Which Python library is best suited for extracting only the main article content from these HTML pages?
Show answer details
Correct answer: C
Beautiful Soup is a powerful Python library designed for parsing HTML and XML documents. It excels at navigating, searching, and modifying the parse tree, making it ideal for extracting specific content (like the main article text within
ortags) while ignoring irrelevant boilerplate content (like menus intags or ads). Whilerequestsis used to fetch the HTML,beautifulsoup4is the tool for cleaning and extracting the desired content. - 10
A development team is using Inference Tables to monitor a deployed RAG application. They notice a sudden spike in requests that result in responses containing fallback messages like 'I cannot answer this question based on the provided information.' This indicates a problem with the retrieval step. Which TWO metrics, available through Inference Tables and the associated monitoring dashboards, would be most direct in diagnosing this retrieval failure? (Choose two.)
Show answer details
Correct answer: A, C
