Skip to main content

Missing Values in SAS vs R: The Differences That Change Your Numbers

sas-to-r
r-programming
clinical
SAS has numeric missing (.) and 27 special missings that sort low and compare as the smallest value; R has NA that propagates and NULL that is not a value. How each behaviour affects filters, sorts, counts and derivations in a SAS-to-R migration, with runnable examples.
Author

Rverse Analytics

Published

October 8, 2026

  • In SAS, numeric missing is less than every number: if AGE < 65 keeps missing ages. In R, AGE < 65 is NA for missing values and subset()/filter() drop them. This single rule explains most record-count differences in a parallel run.
  • SAS has special missings (.A to .Z, ._) used for “not done”, “not applicable” and the like. R has one NA per type; model the reason in a second variable.
  • SAS character missing is a blank string that sorts first; R distinguishes "" from NA_character_.
  • R functions default to propagating NA (mean() returns NA) where SAS procedures drop it. Always state na.rm explicitly in migration code.

Every SAS programmer knows that a missing value is “the smallest number”. Every R programmer knows that NA means “unknown” and that comparisons with it are unknown too. Both are internally consistent, and they produce different record counts the moment a filter touches a variable with missing values. In a migration this appears on day one, so it is worth being precise about it.

Comparisons and filters

data young;
  set adsl;
  if AGE < 65;   /* missing AGE is kept: . < 65 is TRUE in SAS */
run;
adsl <- data.frame(USUBJID = sprintf("01-%03d", 1:6), AGE = c(54, NA, 71, 63, NA, 68))
adsl$AGE < 65
[1]  TRUE    NA FALSE  TRUE    NA FALSE
subset(adsl, AGE < 65)          # R drops the NA rows
  USUBJID AGE
1  01-001  54
4  01-004  63

To reproduce the SAS behaviour explicitly:

subset(adsl, is.na(AGE) | AGE < 65)
  USUBJID AGE
1  01-001  54
2  01-002  NA
4  01-004  63
5  01-005  NA

The SAS-safe idiom is if AGE < 65 and not missing(AGE), and good SAS code has it. If the legacy code does not, the R translation will have fewer records and the equivalence review has to decide which behaviour the specification intended. In our experience the specification nearly always intends the R behaviour, and the SAS program had a latent defect.

Sorting

SAS sorts numeric missing first (ascending) and character blanks first. R’s order() and dplyr::arrange() put NA last by default.

x <- c(3, NA, 1, 2)
sort(x)                      # NA dropped entirely
[1] 1 2 3
sort(x, na.last = TRUE)      # NA last
[1]  1  2  3 NA
sort(x, na.last = FALSE)     # NA first, like SAS
[1] NA  1  2  3

For first./last. logic or “take the earliest visit” derivations, this changes which record is picked. Match SAS during parallel runs with na.last = FALSE or arrange(desc(is.na(x)), x), and write the intended rule into the specification.

Counts and denominators

sex <- c("F", "M", NA, "F", "F", NA)
table(sex)                      # NA excluded, like PROC FREQ default
sex
F M 
3 1 
table(sex, useNA = "ifany")     # like PROC FREQ / MISSING
sex
   F    M <NA> 
   3    1    2 
length(sex)                     # total records
[1] 6
sum(!is.na(sex))                # PROC MEANS N
[1] 4

dplyr::n() counts rows including NA; PROC MEANS N counts non-missing. Use sum(!is.na(x)) wherever a SAS N is reproduced. See PROC FREQ and PROC MEANS in R for the full set of translations.

Aggregation functions

v <- c(2.5, NA, 4.0)
mean(v)
[1] NA
mean(v, na.rm = TRUE)
[1] 3.25

SAS MEAN() across variables in a DATA step and PROC MEANS both skip missing silently. R propagates. The migration standard should require na.rm = TRUE (or the opposite) to be written at every call site, so the choice is visible to the QC reviewer.

Special missing values

SAS numeric missing can be ., ._ or .A through .Z, and they sort in that order, all below every number. Clinical datasets use them for “not done”, “not evaluable”, “below limit of quantification” and so on. R has no equivalent. The clean translation is one value variable plus one reason variable:

lb <- data.frame(USUBJID = c("01", "02", "03", "04"), AVAL = c(5.2, NA, NA, 7.9),
                 AVALC = c("5.2", "<LLOQ", "ND", "7.9"))
lb$missing_reason <- ifelse(is.na(lb$AVAL), lb$AVALC, NA_character_)
lb
  USUBJID AVAL AVALC missing_reason
1      01  5.2   5.2           <NA>
2      02   NA <LLOQ          <LLOQ
3      03   NA    ND             ND
4      04  7.9   7.9           <NA>

In SDTM/ADaM the analogous information lives in --STAT, --REASND and AVALC, so this is also the standards-compliant form. haven::read_sas() and haven::read_xpt() preserve SAS special missings as tagged NA, which haven::is_tagged_na() and haven::na_tag() can read back if you need the original codes during parallel runs.

Character missing

s <- c("A", "", NA)
is.na(s)
[1] FALSE FALSE  TRUE
s == ""
[1] FALSE  TRUE    NA
nchar(s)
[1]  1  0 NA

SAS has only the blank. R distinguishes empty string from NA, and haven imports SAS blanks as "", not NA. Decide once: our standard converts imported "" to NA_character_ at the boundary (dplyr::na_if(x, "")), so downstream code has one notion of missing.

Arithmetic and derivations

SAS: any arithmetic with a missing operand is missing, and the log notes “Missing values were generated”. R: the same, silently. The difference is the SUM() function versus the + operator in SAS: SUM(a, b) ignores missing, a + b propagates. In R sum(a, b, na.rm = TRUE) is the first and a + b the second. A macro that used SUM() for a total score must be translated with rowSums(..., na.rm = TRUE) or the scores will differ for every subject with one missing item. Check the scoring rule in the specification: often a minimum number of non-missing items is required before a total is computed, and that rule belongs in the function.

items <- data.frame(q1 = c(3, 4, NA), q2 = c(2, NA, NA), q3 = c(5, 4, 3))
items$total_sas_sum  <- rowSums(items[, 1:3], na.rm = TRUE)          # SUM(q1,q2,q3)
items$total_plus     <- items$q1 + items$q2 + items$q3               # q1+q2+q3
items$n_items        <- rowSums(!is.na(items[, 1:3]))
items$total_rule     <- ifelse(items$n_items >= 2, items$total_sas_sum * 3 / items$n_items, NA)  # prorated, ≥2 items
items
  q1 q2 q3 total_sas_sum total_plus n_items total_rule
1  3  2  5            10         10       3         10
2  4 NA  4             8         NA       2         12
3 NA NA  3             3         NA       1         NA

A checklist for migration code review

  1. Every comparison on a variable that can be missing states what happens to NA.
  2. Every sort/arrange on such a variable states na.last.
  3. Every aggregation states na.rm.
  4. "" is converted to NA at import, once.
  5. Special missings are carried as a reason variable, not lost.
  6. Record counts are logged before and after each filter and compared to the SAS log.

These six lines prevent the majority of “R gives a different N” tickets. The rest are covered by the numerical tolerance rules and the rounding post.