Skip to content

HPE0-V30 HPE AI Fundamentals Practice Questions

Prepare for HPE0-V30 with more than an answer.

150 questions in the full set12 sample questionsUpdated Jul 8, 2026
Exam fee
$91 USD
Time limit
90 minutes
Questions on the exam
60
Passing score
65%
Level
HPE ATP - AI solutions (Associate/Technical Professional)
Domains covered on the exam 3
  1. Introduction To GenAI and Industry Specific Applications44%
  2. Data Manipulation-Cleaning and Labelling44%
  3. NVIDIA Concepts12%
  1. 1

    When adapting the Transformer architecture from NLP to Computer Vision tasks, how does the Vision Transformer (ViT) architecture initially format a 2D image so it can be processed by standard self-attention layers?

    Show answer details

    Correct answer: C

    The core innovation of the Vision Transformer (ViT) is treating image patches like words (tokens) in a sentence. The image is divided into a grid of fixed-size patches. Each patch is flattened into a 1D vector and then passed through a trainable linear projection layer to create the initial embeddings that the Transformer encoder expects.

  2. 2

    In an encoder-decoder Transformer model applied to image captioning, how does the cross-attention mechanism operate during the generation of the text caption?

    Show answer details

    Correct answer: C

    In cross-attention (or encoder-decoder attention), the model aligns two different sequences. The generating sequence (the text decoder) provides the Queries, asking 'what information do I need next?'. The source sequence (the image encoder) provides the Keys and Values, supplying the actual visual feature information to answer that query.

  3. 3

    What is the primary difference between the 'pre-training' phase and the 'inference' phase of a Large Language Model (LLM)?

    Show answer details

    Correct answer: C

    Pre-training is the highly compute-intensive phase where the model learns language patterns by predicting tokens across massive datasets, constantly updating its internal weights via backpropagation. Inference is the application phase where the model uses those frozen, learned weights to generate responses to new user prompts.

  4. 4

    What fundamental limitation of Recurrent Neural Networks (RNNs) did the Transformer architecture primarily resolve by introducing the self-attention mechanism?

    Show answer details

    Correct answer: D

    RNNs process sequences step-by-step, making parallel computation impossible and causing the vanishing gradient problem over long distances. The Transformer's self-attention mechanism evaluates all tokens simultaneously, allowing massive parallelization and direct connections between distant words.

  5. 5

    During the computation of scaled dot-product attention in a Transformer, the dot product of the Query (Q) and Key (K) matrices is divided by the square root of the dimension of the key (sqrt(d_k)). What is the primary mathematical reason for this scaling factor?

    flowchart LR Q[Query] --> Dot[Dot Product Q*K^T] K[Key] --> Dot Dot --> Scale[Scale by 1/sqrt d_k] Scale --> Softmax[Softmax] Softmax --> Mult[Multiply with Value] V[Value] --> Mult Mult --> Out[Attention Output]
    Show answer details

    Correct answer: C

    As the dimension of the key vectors (d_k) increases, the variance of the dot product increases, leading to very large values. Feeding these large values into a softmax function pushes the outputs to 1 or 0, resulting in extremely small (vanishing) gradients during backpropagation. Scaling by the square root of d_k normalizes the variance.

  6. 6

    Which TWO of the following statements accurately describe the role and implementation of Positional Encoding in the standard Transformer architecture? (Select TWO)

    Show answer details

    Correct answer: B, C

    The original "Attention Is All You Need" paper proposed using continuous sinusoidal functions (sine and cosine) of varying frequencies to inject positional information, allowing the model to easily learn to attend by relative positions.

    Self-attention operations treat the input sequence as a "bag of words." Without positional encoding, the model would not be able to distinguish the order of tokens, which is critical for understanding language syntax and semantics.

Create an account to continue.