Skip to main content

PROC FREQ, PROC MEANS and PROC UNIVARIATE in R: Side-by-Side Translations

sas-to-r
r-programming
clinical
The three SAS procedures behind most clinical summary tables, translated to base R and dplyr with matching output: frequencies and percentages with missing handling, by-group descriptive statistics, percentiles with SAS definitions, and the options that cause mismatches.
Author

Rverse Analytics

Published

October 8, 2026

  • PROC FREQ → table()/dplyr::count(); watch the denominator (MISSING option) and the percent rounding.
  • PROC MEANS → dplyr::summarise() with across(); n must count non-missing values and NWAY corresponds to grouping by all class variables.
  • PROC UNIVARIATE percentiles default to SAS definition 5; R’s quantile() defaults to type 7. Use type = 2 to match SAS PCTLDEF=5.
  • Write each translation once as a function, test it against a SAS export, and reuse it. See from SAS macros to R functions.

Three procedures produce the bulk of a clinical TLF package: PROC FREQ for counts and percentages, PROC MEANS for descriptive statistics, and PROC UNIVARIATE for percentiles and distribution checks. Their R translations are short, but each has one or two options whose defaults differ and will show up in a parallel run. Here they are, side by side.

library(dplyr)
set.seed(3)
adsl <- data.frame(
  USUBJID = sprintf("01-%03d", 1:60),
  TRT01A  = rep(c("Placebo", "Drug"), each = 30),
  SEX     = sample(c("F", "M"), 60, replace = TRUE),
  RACE    = sample(c("WHITE", "BLACK OR AFRICAN AMERICAN", "ASIAN", NA), 60, replace = TRUE, prob = c(.6, .2, .15, .05)),
  AGE     = round(rnorm(60, 59, 10)),
  BMIBL   = round(rnorm(60, 27, 4), 1)
)

PROC FREQ

proc freq data=adsl;
  tables TRT01A*RACE / nocol nopct;
run;

The default excludes missing values from the table and from the percentage denominator. Adding / MISSING includes them in both.

# Default PROC FREQ: missings excluded from counts and denominator
freq_tab <- function(data, by, var, include_missing = FALSE) {
  d <- if (include_missing) data else data[!is.na(data[[var]]), ]
  d |>
    count(.data[[by]], .data[[var]], name = "n") |>
    group_by(.data[[by]]) |>
    mutate(N = sum(n), pct = n / N * 100) |>
    ungroup()
}
freq_tab(adsl, "TRT01A", "RACE")
# A tibble: 6 × 5
  TRT01A  RACE                          n     N   pct
  <chr>   <chr>                     <int> <int> <dbl>
1 Drug    ASIAN                         4    30  13.3
2 Drug    BLACK OR AFRICAN AMERICAN     7    30  23.3
3 Drug    WHITE                        19    30  63.3
4 Placebo ASIAN                         5    28  17.9
5 Placebo BLACK OR AFRICAN AMERICAN     9    28  32.1
6 Placebo WHITE                        14    28  50  
# PROC FREQ ... / MISSING
freq_tab(adsl, "TRT01A", "RACE", include_missing = TRUE)
# A tibble: 7 × 5
  TRT01A  RACE                          n     N   pct
  <chr>   <chr>                     <int> <int> <dbl>
1 Drug    ASIAN                         4    30 13.3 
2 Drug    BLACK OR AFRICAN AMERICAN     7    30 23.3 
3 Drug    WHITE                        19    30 63.3 
4 Placebo ASIAN                         5    30 16.7 
5 Placebo BLACK OR AFRICAN AMERICAN     9    30 30   
6 Placebo WHITE                        14    30 46.7 
7 Placebo <NA>                          2    30  6.67

In a demography table the usual clinical convention is a third variant: the denominator is the number of subjects in the arm (from ADSL), missing is shown as its own row, and the percentage for missing is often suppressed. Encode that convention in the function, not in each program.

Mismatch sources: the denominator choice above; row-percent vs column-percent (PROC FREQ prints both unless suppressed); and the display rounding of percentages, covered in rounding differences between SAS and R.

PROC MEANS

proc means data=adsl n mean std median min max maxdec=1 nway;
  class TRT01A;
  var AGE BMIBL;
run;
adsl |>
  group_by(TRT01A) |>
  summarise(across(c(AGE, BMIBL), list(
    n      = ~ sum(!is.na(.x)),
    mean   = ~ mean(.x, na.rm = TRUE),
    sd     = ~ sd(.x, na.rm = TRUE),
    median = ~ median(.x, na.rm = TRUE),
    min    = ~ min(.x, na.rm = TRUE),
    max    = ~ max(.x, na.rm = TRUE)
  ), .names = "{.col}_{.fn}"), .groups = "drop")
# A tibble: 2 × 13
  TRT01A  AGE_n AGE_mean AGE_sd AGE_median AGE_min AGE_max BMIBL_n BMIBL_mean
  <chr>   <int>    <dbl>  <dbl>      <dbl>   <dbl>   <dbl>   <int>      <dbl>
1 Drug       30     58.9   7.60       58        40      75      30       28.3
2 Placebo    30     61.8   7.35       62.5      46      76      30       26.0
# ℹ 4 more variables: BMIBL_sd <dbl>, BMIBL_median <dbl>, BMIBL_min <dbl>,
#   BMIBL_max <dbl>

Mismatch sources:

  • n in PROC MEANS is the count of non-missing values; dplyr::n() counts rows. Use sum(!is.na(x)).
  • NWAY means “only the highest-level combination of class variables”. Grouping by all class variables in R gives exactly that. Without NWAY, SAS also outputs the marginal totals (_TYPE_ < max), which in R you add with a second summarise over fewer groups or add_overall()-style helpers.
  • MAXDEC= is display only. Do not round before summarising.
  • SAS STD and R sd() both use the n−1 denominator; VARDEF= changes that in SAS and must be matched if used.

PROC UNIVARIATE and percentiles

proc univariate data=adsl;
  class TRT01A;
  var AGE;
  output out=pct pctlpts=25 50 75 pctlpre=P;
run;

SAS’s default percentile definition is PCTLDEF=5 (empirical distribution function with averaging). R’s quantile() default is type 7 (linear interpolation), and the two differ for small groups.

x <- adsl$AGE[adsl$TRT01A == "Drug"]
rbind(
  r_default_type7 = quantile(x, c(.25, .5, .75), type = 7),
  sas_pctldef5    = quantile(x, c(.25, .5, .75), type = 2)
)
                25% 50%   75%
r_default_type7  54  58 64.75
sas_pctldef5     54  58 65.00

Hyndman and Fan’s type 2 is the SAS default (PCTLDEF=5); type 6 corresponds to PCTLDEF=4 and type 3 to PCTLDEF=2. Put the chosen type in the standards document and in the function signature:

pctl <- function(x, probs = c(.25, .5, .75), sas_def = 5) {
  type <- c(`1` = 3, `2` = 3, `3` = 1, `4` = 6, `5` = 2)[as.character(sas_def)]
  quantile(x, probs, type = type, na.rm = TRUE, names = FALSE)
}
pctl(x)
[1] 54 58 65

Medians agree between type 2 and type 7 for odd n and usually for even n; quartiles are where differences appear. The median of a Q1-Q3 display in a demography table is therefore safe, but the quartiles themselves need the matching type.

Other PROC UNIVARIATE outputs: skewness and kurtosis use the sample-adjusted formulas in SAS (VARDEF=DF); e1071::skewness(type = 2) and kurtosis(type = 2) match them. Normality tests: Shapiro-Wilk is shapiro.test(); SAS reports it only for n ≤ 2000, R up to 5000.

A translation table to keep

SAS R Watch for
PROC FREQ count() + mutate(pct = n/sum(n)*100) Denominator and MISSING option
PROC FREQ ... / CHISQ chisq.test(correct = FALSE) SAS does not apply Yates’ correction by default
PROC FREQ ... / FISHER fisher.test() Two-sided p agrees; one-sided definitions differ
PROC MEANS summarise(across(...)) n = non-missing; NWAY vs marginals
PROC UNIVARIATE percentiles quantile(type = 2) PCTLDEF default 5
PROC UNIVARIATE ... NORMAL shapiro.test() Sample-size limits
PROC SORT NODUPKEY distinct(..., .keep_all = TRUE) Keeps first occurrence in both
PROC TRANSPOSE tidyr::pivot_wider()/pivot_longer() Variable naming of transposed columns

Each row of this table is one function in a migration package and one test against a SAS export. That is the whole method: translate once, test once, reuse everywhere. For the full programme see the five-phase migration roadmap and the PROC-to-R equivalence map.