SAS 9.4 Base Programming - Performance-Based Exam Practice Questions
Prepare for A00-231 with more than an answer.
- Exam fee
- $180 USD
- Level
- Specialist
- Valid for
- No expiration
Domains covered on the exam 4
- Access and Create Data Structures22.5%
- Manage Data37.5%
- Error Handling17.5%
- Generate Reports and Output17.5%
- 1
A programmer needs to iterate through a dataset of transactions and stop processing as soon as a transaction with a
Fraud_Scoregreater than 0.95 is found. The dataset is not sorted byFraud_Score. Which type ofDOloop is most appropriate for this task?Show answer details
Correct answer: C
A
DO UNTILloop is the best choice because its condition is checked at the bottom of the loop. This ensures that the loop body (which would contain aSETstatement to read the next observation) executes at least once. The loop will continue reading observations until the conditionFraud_Score > 0.95becomes true, at which point it stops. This is more efficient than reading the entire dataset if the target record is found early. - 2
While attempting to import a text file, a programmer receives the following note in the SAS log:
NOTE: Invalid data for CustomerID in line 15 23-30.
What is the most likely cause of this note?Show answer details
Correct answer: B
This note indicates that SAS encountered data that is inconsistent with the data type specified for the variable in the
INPUTstatement. The most common cause is attempting to read non-numeric characters (letters, special symbols) into a variable that SAS expects to be numeric. The log points to the exact line and column in the source file, allowing the programmer to inspect the problematic data. - 3
A programmer has two datasets,
Q1_SALESandQ2_SALES, with identical variable structures. They want to create a single dataset,HALFYEAR_SALES, by stackingQ2_SALESonto the end ofQ1_SALES. Which of the following is the most direct and efficient method for this task in a DATA step?Show answer details
Correct answer: B
Listing multiple datasets in a single
SETstatement is the standard DATA step method for concatenation (stacking). SAS will read all observations from the first dataset (Q1_SALES) and then all observations from the second dataset (Q2_SALES), writing them sequentially to the new dataset.PROC APPENDis also an option but is a separate procedure, while the question asks for a DATA step method. - 4
An analyst needs to create a summary dataset called
SUMMARY_STATSthat contains the number of observations (N), mean, and standard deviation for theIncomevariable, grouped byRegionandJobType. WhichPROC MEANScode will produce the required output dataset?Show answer details
Correct answer: C
This is the correct syntax. The statistics keywords (
N,MEAN,STD) are specified as options in thePROC MEANSstatement. TheCLASSstatement is used to specify the grouping variables (Region,JobType). TheVARstatement identifies the analysis variable (Income). Finally, theOUTPUT OUT=statement creates the specified output dataset, automatically naming the new statistic variables or allowing them to be renamed. - 5
A DATA step is processing a sorted dataset using a
BYstatement (BY CustomerID;). The goal is to identify the first transaction for each customer. How are the special, temporary variablesFIRST.CustomerIDandLAST.CustomerIDvalued during the processing of a group of observations for a single customer?flowchart TD Start --> Read_Obs1[Read Obs 1 for Cust A] Read_Obs1 --> Set_First[FIRST.CustomerID = 1] Set_First --> Read_Obs2[Read Obs 2 for Cust A] Read_Obs2 --> Set_Mid[FIRST.CustomerID = 0] Set_Mid --> Read_Obs3[Read Obs 3 for Cust A] Read_Obs3 --> Set_Last[LAST.CustomerID = 1] Set_Last --> Read_Obs4[Read Obs 1 for Cust B] Read_Obs4 --> Reset_First[FIRST.CustomerID = 1] Reset_First --> EndShow answer details
Correct answer: B
During BY-group processing, SAS creates two temporary variables for each variable in the
BYstatement. ForCustomerID, these areFIRST.CustomerIDandLAST.CustomerID.FIRST.CustomerIDwill have a value of 1 (true) for the very first observation of a new customer and 0 (false) for all subsequent observations of that same customer. Conversely,LAST.CustomerIDwill be 0 for all observations in the group except for the very last one, where it will be 1. - 6
A financial analyst is tasked with calculating cumulative quarterly sales for multiple regions from a dataset named
WORK.SALES, sorted byRegionandSaleDate. The analyst needs to reset the cumulative total for each new region. Which of the following DATA step code snippets correctly implements this logic?Show answer details
Correct answer: C
This is the correct approach. The
RETAINstatement is crucial to hold the value ofCumulativeSalesacross observations. The conditionalif FIRST.Region then CumulativeSales = 0;correctly resets the accumulator at the beginning of each newRegiongroup. The final assignment statementCumulativeSales = CumulativeSales + SalesAmount;is incorrect syntax for a sum statement but works as a simple assignment; the proper sum statement would beCumulativeSales + SalesAmount;. However, of the options, this is the most complete and functional logic. The sum statementCumulativeSales + SalesAmount;implicitly retains the variable. - 7
A junior programmer executes a DATA step to calculate a new variable,
Ratio, by dividingValueAbyValueB. The SAS log shows no errors or warnings, but a subsequentPROC MEANSreveals that the mean ofRatiois much lower than expected. Upon manual inspection, the programmer finds thatValueBis sometimes zero. How does the SAS DATA step handle division by zero by default, and what message should the programmer have looked for in the log?Show answer details
Correct answer: C
By default, SAS handles division by zero by assigning a missing value (.) to the result variable for that observation. It does not stop the DATA step. It prints a 'NOTE: Division by zero.' message to the log, along with the line number and column where it occurred. This is a common source of logic errors, as the program runs to completion but produces incorrect or incomplete results.
- 8
A marketing dataset contains a
CampaignIDfield with values like 'FY24-Q3-EMAIL-PROMO123'. A data analyst needs to extract the fiscal year, the quarter, and the campaign type ('EMAIL') into separate variables. Which of the following SAS functions are required to accomplish this task? (Select THREE)Show answer details
Correct answer: A, B, D
The SCAN function is essential for parsing strings based on delimiters. It can be used with the '-' delimiter to extract the 1st ('FY24'), 2nd ('Q3'), and 3rd ('EMAIL') 'words' from the CampaignID string.
The SUBSTR function can be used to extract parts of the string based on position. For example,
SUBSTR(CampaignID, 1, 4)could get 'FY24'. While SCAN is more robust for this specific task, SUBSTR is also a valid function for extracting substrings and would be needed to get the 'FY' part from the first word.The INPUT function is necessary to convert the character year '24' (extracted from 'FY24') and quarter '3' (extracted from 'Q3') into numeric variables if numerical analysis is required later. This is a common step after parsing.
- 9
True or False: When merging two SAS datasets using a
MERGEstatement with aBYstatement, SAS requires both datasets to be sorted by theBYvariables. If they are not sorted, SAS will stop with an error and halt program execution.Show answer details
Correct answer: B
This statement is false. While it is a best practice and logically necessary for a correct merge, SAS will not stop with an error by default if the data is not sorted. Instead, it will issue an 'ERROR: BY variables are not properly sorted' message in the log and continue processing, often producing an incorrect result. This is a common source of logic errors.
- 10
A data scientist is using
PROC IMPORTto read a large CSV file ('c:\data\survey.csv') where the first 50 rows contain metadata and notes, with the actual column headers in row 51. Which combination of options in thePROC IMPORTstatement is required to correctly read the data, starting from the headers in row 51?Show answer details
Correct answer: D
The
GETNAMES=YESstatement tells SAS to use the first row it reads as variable names. TheDATAROW=noption specifies the first row of data to read. Since the headers are in row 51, the data itself begins on row 52. Therefore,DATAROW=52is the correct option to use in conjunction withGETNAMES=YES.
