adsl <- data.frame(USUBJID = sprintf("01-%03d", 1:6), AGE = c(54, NA, 71, 63, NA, 68))
adsl$AGE < 65[1] TRUE NA FALSE TRUE NA FALSE
subset(adsl, AGE < 65) # R drops the NA rows USUBJID AGE
1 01-001 54
4 01-004 63
Rverse Analytics
October 8, 2026
if AGE < 65 keeps missing ages. In R, AGE < 65 is NA for missing values and subset()/filter() drop them. This single rule explains most record-count differences in a parallel run..A to .Z, ._) used for “not done”, “not applicable” and the like. R has one NA per type; model the reason in a second variable."" from NA_character_.mean() returns NA) where SAS procedures drop it. Always state na.rm explicitly in migration code.Every SAS programmer knows that a missing value is “the smallest number”. Every R programmer knows that NA means “unknown” and that comparisons with it are unknown too. Both are internally consistent, and they produce different record counts the moment a filter touches a variable with missing values. In a migration this appears on day one, so it is worth being precise about it.
data young;
set adsl;
if AGE < 65; /* missing AGE is kept: . < 65 is TRUE in SAS */
run;
[1] TRUE NA FALSE TRUE NA FALSE
USUBJID AGE
1 01-001 54
4 01-004 63
To reproduce the SAS behaviour explicitly:
The SAS-safe idiom is if AGE < 65 and not missing(AGE), and good SAS code has it. If the legacy code does not, the R translation will have fewer records and the equivalence review has to decide which behaviour the specification intended. In our experience the specification nearly always intends the R behaviour, and the SAS program had a latent defect.
SAS sorts numeric missing first (ascending) and character blanks first. R’s order() and dplyr::arrange() put NA last by default.
[1] 1 2 3
[1] 1 2 3 NA
[1] NA 1 2 3
For first./last. logic or “take the earliest visit” derivations, this changes which record is picked. Match SAS during parallel runs with na.last = FALSE or arrange(desc(is.na(x)), x), and write the intended rule into the specification.
sex
F M
3 1
sex
F M <NA>
3 1 2
[1] 6
[1] 4
dplyr::n() counts rows including NA; PROC MEANS N counts non-missing. Use sum(!is.na(x)) wherever a SAS N is reproduced. See PROC FREQ and PROC MEANS in R for the full set of translations.
SAS MEAN() across variables in a DATA step and PROC MEANS both skip missing silently. R propagates. The migration standard should require na.rm = TRUE (or the opposite) to be written at every call site, so the choice is visible to the QC reviewer.
SAS numeric missing can be ., ._ or .A through .Z, and they sort in that order, all below every number. Clinical datasets use them for “not done”, “not evaluable”, “below limit of quantification” and so on. R has no equivalent. The clean translation is one value variable plus one reason variable:
USUBJID AVAL AVALC missing_reason
1 01 5.2 5.2 <NA>
2 02 NA <LLOQ <LLOQ
3 03 NA ND ND
4 04 7.9 7.9 <NA>
In SDTM/ADaM the analogous information lives in --STAT, --REASND and AVALC, so this is also the standards-compliant form. haven::read_sas() and haven::read_xpt() preserve SAS special missings as tagged NA, which haven::is_tagged_na() and haven::na_tag() can read back if you need the original codes during parallel runs.
SAS has only the blank. R distinguishes empty string from NA, and haven imports SAS blanks as "", not NA. Decide once: our standard converts imported "" to NA_character_ at the boundary (dplyr::na_if(x, "")), so downstream code has one notion of missing.
SAS: any arithmetic with a missing operand is missing, and the log notes “Missing values were generated”. R: the same, silently. The difference is the SUM() function versus the + operator in SAS: SUM(a, b) ignores missing, a + b propagates. In R sum(a, b, na.rm = TRUE) is the first and a + b the second. A macro that used SUM() for a total score must be translated with rowSums(..., na.rm = TRUE) or the scores will differ for every subject with one missing item. Check the scoring rule in the specification: often a minimum number of non-missing items is required before a total is computed, and that rule belongs in the function.
items <- data.frame(q1 = c(3, 4, NA), q2 = c(2, NA, NA), q3 = c(5, 4, 3))
items$total_sas_sum <- rowSums(items[, 1:3], na.rm = TRUE) # SUM(q1,q2,q3)
items$total_plus <- items$q1 + items$q2 + items$q3 # q1+q2+q3
items$n_items <- rowSums(!is.na(items[, 1:3]))
items$total_rule <- ifelse(items$n_items >= 2, items$total_sas_sum * 3 / items$n_items, NA) # prorated, ≥2 items
items q1 q2 q3 total_sas_sum total_plus n_items total_rule
1 3 2 5 10 10 3 10
2 4 NA 4 8 NA 2 12
3 NA NA 3 3 NA 1 NA
sort/arrange on such a variable states na.last.na.rm."" is converted to NA at import, once.These six lines prevent the majority of “R gives a different N” tickets. The rest are covered by the numerical tolerance rules and the rounding post.