Skip to content

SAS 9.4 Base Programming - Performance-Based Exam Practice Questions

Prepare for A00-231 with more than an answer.

217 questions in the full set20 sample questionsUpdated Oct 18, 2025
Exam fee
$180 USD
Level
Specialist
Valid for
No expiration
Domains covered on the exam 4
  1. Access and Create Data Structures22.5%
  2. Manage Data37.5%
  3. Error Handling17.5%
  4. Generate Reports and Output17.5%
  1. 1

    A programmer needs to iterate through a dataset of transactions and stop processing as soon as a transaction with a Fraud_Score greater than 0.95 is found. The dataset is not sorted by Fraud_Score. Which type of DO loop is most appropriate for this task?

    Show answer details

    Correct answer: C

    A DO UNTIL loop is the best choice because its condition is checked at the bottom of the loop. This ensures that the loop body (which would contain a SET statement to read the next observation) executes at least once. The loop will continue reading observations until the condition Fraud_Score > 0.95 becomes true, at which point it stops. This is more efficient than reading the entire dataset if the target record is found early.

  2. 2

    While attempting to import a text file, a programmer receives the following note in the SAS log:
    NOTE: Invalid data for CustomerID in line 15 23-30.
    What is the most likely cause of this note?

    Show answer details

    Correct answer: B

    This note indicates that SAS encountered data that is inconsistent with the data type specified for the variable in the INPUT statement. The most common cause is attempting to read non-numeric characters (letters, special symbols) into a variable that SAS expects to be numeric. The log points to the exact line and column in the source file, allowing the programmer to inspect the problematic data.

  3. 3

    A programmer has two datasets, Q1_SALES and Q2_SALES, with identical variable structures. They want to create a single dataset, HALFYEAR_SALES, by stacking Q2_SALES onto the end of Q1_SALES. Which of the following is the most direct and efficient method for this task in a DATA step?

    Show answer details

    Correct answer: B

    Listing multiple datasets in a single SET statement is the standard DATA step method for concatenation (stacking). SAS will read all observations from the first dataset (Q1_SALES) and then all observations from the second dataset (Q2_SALES), writing them sequentially to the new dataset. PROC APPEND is also an option but is a separate procedure, while the question asks for a DATA step method.

  4. 4

    An analyst needs to create a summary dataset called SUMMARY_STATS that contains the number of observations (N), mean, and standard deviation for the Income variable, grouped by Region and JobType. Which PROC MEANS code will produce the required output dataset?

    Show answer details

    Correct answer: C

    This is the correct syntax. The statistics keywords (N, MEAN, STD) are specified as options in the PROC MEANS statement. The CLASS statement is used to specify the grouping variables (Region, JobType). The VAR statement identifies the analysis variable (Income). Finally, the OUTPUT OUT= statement creates the specified output dataset, automatically naming the new statistic variables or allowing them to be renamed.

  5. 5

    A DATA step is processing a sorted dataset using a BY statement (BY CustomerID;). The goal is to identify the first transaction for each customer. How are the special, temporary variables FIRST.CustomerID and LAST.CustomerID valued during the processing of a group of observations for a single customer?

    flowchart TD Start --> Read_Obs1[Read Obs 1 for Cust A] Read_Obs1 --> Set_First[FIRST.CustomerID = 1] Set_First --> Read_Obs2[Read Obs 2 for Cust A] Read_Obs2 --> Set_Mid[FIRST.CustomerID = 0] Set_Mid --> Read_Obs3[Read Obs 3 for Cust A] Read_Obs3 --> Set_Last[LAST.CustomerID = 1] Set_Last --> Read_Obs4[Read Obs 1 for Cust B] Read_Obs4 --> Reset_First[FIRST.CustomerID = 1] Reset_First --> End
    Show answer details

    Correct answer: B

    During BY-group processing, SAS creates two temporary variables for each variable in the BY statement. For CustomerID, these are FIRST.CustomerID and LAST.CustomerID. FIRST.CustomerID will have a value of 1 (true) for the very first observation of a new customer and 0 (false) for all subsequent observations of that same customer. Conversely, LAST.CustomerID will be 0 for all observations in the group except for the very last one, where it will be 1.

  6. 6

    A financial analyst is tasked with calculating cumulative quarterly sales for multiple regions from a dataset named WORK.SALES, sorted by Region and SaleDate. The analyst needs to reset the cumulative total for each new region. Which of the following DATA step code snippets correctly implements this logic?

    Show answer details

    Correct answer: C

    This is the correct approach. The RETAIN statement is crucial to hold the value of CumulativeSales across observations. The conditional if FIRST.Region then CumulativeSales = 0; correctly resets the accumulator at the beginning of each new Region group. The final assignment statement CumulativeSales = CumulativeSales + SalesAmount; is incorrect syntax for a sum statement but works as a simple assignment; the proper sum statement would be CumulativeSales + SalesAmount;. However, of the options, this is the most complete and functional logic. The sum statement CumulativeSales + SalesAmount; implicitly retains the variable.

  7. 7

    A junior programmer executes a DATA step to calculate a new variable, Ratio, by dividing ValueA by ValueB. The SAS log shows no errors or warnings, but a subsequent PROC MEANS reveals that the mean of Ratio is much lower than expected. Upon manual inspection, the programmer finds that ValueB is sometimes zero. How does the SAS DATA step handle division by zero by default, and what message should the programmer have looked for in the log?

    Show answer details

    Correct answer: C

    By default, SAS handles division by zero by assigning a missing value (.) to the result variable for that observation. It does not stop the DATA step. It prints a 'NOTE: Division by zero.' message to the log, along with the line number and column where it occurred. This is a common source of logic errors, as the program runs to completion but produces incorrect or incomplete results.

  8. 8

    A marketing dataset contains a CampaignID field with values like 'FY24-Q3-EMAIL-PROMO123'. A data analyst needs to extract the fiscal year, the quarter, and the campaign type ('EMAIL') into separate variables. Which of the following SAS functions are required to accomplish this task? (Select THREE)

    Show answer details

    Correct answer: A, B, D

    The SCAN function is essential for parsing strings based on delimiters. It can be used with the '-' delimiter to extract the 1st ('FY24'), 2nd ('Q3'), and 3rd ('EMAIL') 'words' from the CampaignID string.

    The SUBSTR function can be used to extract parts of the string based on position. For example, SUBSTR(CampaignID, 1, 4) could get 'FY24'. While SCAN is more robust for this specific task, SUBSTR is also a valid function for extracting substrings and would be needed to get the 'FY' part from the first word.

    The INPUT function is necessary to convert the character year '24' (extracted from 'FY24') and quarter '3' (extracted from 'Q3') into numeric variables if numerical analysis is required later. This is a common step after parsing.

  9. 9

    True or False: When merging two SAS datasets using a MERGE statement with a BY statement, SAS requires both datasets to be sorted by the BY variables. If they are not sorted, SAS will stop with an error and halt program execution.

    Show answer details

    Correct answer: B

    This statement is false. While it is a best practice and logically necessary for a correct merge, SAS will not stop with an error by default if the data is not sorted. Instead, it will issue an 'ERROR: BY variables are not properly sorted' message in the log and continue processing, often producing an incorrect result. This is a common source of logic errors.

  10. 10

    A data scientist is using PROC IMPORT to read a large CSV file ('c:\data\survey.csv') where the first 50 rows contain metadata and notes, with the actual column headers in row 51. Which combination of options in the PROC IMPORT statement is required to correctly read the data, starting from the headers in row 51?

    Show answer details

    Correct answer: D

    The GETNAMES=YES statement tells SAS to use the first row it reads as variable names. The DATAROW=n option specifies the first row of data to read. Since the headers are in row 51, the data itself begins on row 52. Therefore, DATAROW=52 is the correct option to use in conjunction with GETNAMES=YES.

Create an account to continue.