library(dplyr)
set.seed(3)
adsl <- data.frame(
USUBJID = sprintf("01-%03d", 1:60),
TRT01A = rep(c("Placebo", "Drug"), each = 30),
SEX = sample(c("F", "M"), 60, replace = TRUE),
RACE = sample(c("WHITE", "BLACK OR AFRICAN AMERICAN", "ASIAN", NA), 60, replace = TRUE, prob = c(.6, .2, .15, .05)),
AGE = round(rnorm(60, 59, 10)),
BMIBL = round(rnorm(60, 27, 4), 1)
)PROC FREQ, PROC MEANS and PROC UNIVARIATE in R: Side-by-Side Translations
PROC FREQ→table()/dplyr::count(); watch the denominator (MISSINGoption) and the percent rounding.PROC MEANS→dplyr::summarise()withacross();nmust count non-missing values andNWAYcorresponds to grouping by all class variables.PROC UNIVARIATEpercentiles default to SAS definition 5; R’squantile()defaults to type 7. Usetype = 2to match SAS PCTLDEF=5.- Write each translation once as a function, test it against a SAS export, and reuse it. See from SAS macros to R functions.
Three procedures produce the bulk of a clinical TLF package: PROC FREQ for counts and percentages, PROC MEANS for descriptive statistics, and PROC UNIVARIATE for percentiles and distribution checks. Their R translations are short, but each has one or two options whose defaults differ and will show up in a parallel run. Here they are, side by side.
PROC FREQ
proc freq data=adsl;
tables TRT01A*RACE / nocol nopct;
run;
The default excludes missing values from the table and from the percentage denominator. Adding / MISSING includes them in both.
# Default PROC FREQ: missings excluded from counts and denominator
freq_tab <- function(data, by, var, include_missing = FALSE) {
d <- if (include_missing) data else data[!is.na(data[[var]]), ]
d |>
count(.data[[by]], .data[[var]], name = "n") |>
group_by(.data[[by]]) |>
mutate(N = sum(n), pct = n / N * 100) |>
ungroup()
}
freq_tab(adsl, "TRT01A", "RACE")# A tibble: 6 × 5
TRT01A RACE n N pct
<chr> <chr> <int> <int> <dbl>
1 Drug ASIAN 4 30 13.3
2 Drug BLACK OR AFRICAN AMERICAN 7 30 23.3
3 Drug WHITE 19 30 63.3
4 Placebo ASIAN 5 28 17.9
5 Placebo BLACK OR AFRICAN AMERICAN 9 28 32.1
6 Placebo WHITE 14 28 50
# PROC FREQ ... / MISSING
freq_tab(adsl, "TRT01A", "RACE", include_missing = TRUE)# A tibble: 7 × 5
TRT01A RACE n N pct
<chr> <chr> <int> <int> <dbl>
1 Drug ASIAN 4 30 13.3
2 Drug BLACK OR AFRICAN AMERICAN 7 30 23.3
3 Drug WHITE 19 30 63.3
4 Placebo ASIAN 5 30 16.7
5 Placebo BLACK OR AFRICAN AMERICAN 9 30 30
6 Placebo WHITE 14 30 46.7
7 Placebo <NA> 2 30 6.67
In a demography table the usual clinical convention is a third variant: the denominator is the number of subjects in the arm (from ADSL), missing is shown as its own row, and the percentage for missing is often suppressed. Encode that convention in the function, not in each program.
Mismatch sources: the denominator choice above; row-percent vs column-percent (PROC FREQ prints both unless suppressed); and the display rounding of percentages, covered in rounding differences between SAS and R.
PROC MEANS
proc means data=adsl n mean std median min max maxdec=1 nway;
class TRT01A;
var AGE BMIBL;
run;
adsl |>
group_by(TRT01A) |>
summarise(across(c(AGE, BMIBL), list(
n = ~ sum(!is.na(.x)),
mean = ~ mean(.x, na.rm = TRUE),
sd = ~ sd(.x, na.rm = TRUE),
median = ~ median(.x, na.rm = TRUE),
min = ~ min(.x, na.rm = TRUE),
max = ~ max(.x, na.rm = TRUE)
), .names = "{.col}_{.fn}"), .groups = "drop")# A tibble: 2 × 13
TRT01A AGE_n AGE_mean AGE_sd AGE_median AGE_min AGE_max BMIBL_n BMIBL_mean
<chr> <int> <dbl> <dbl> <dbl> <dbl> <dbl> <int> <dbl>
1 Drug 30 58.9 7.60 58 40 75 30 28.3
2 Placebo 30 61.8 7.35 62.5 46 76 30 26.0
# ℹ 4 more variables: BMIBL_sd <dbl>, BMIBL_median <dbl>, BMIBL_min <dbl>,
# BMIBL_max <dbl>
Mismatch sources:
nin PROC MEANS is the count of non-missing values;dplyr::n()counts rows. Usesum(!is.na(x)).NWAYmeans “only the highest-level combination of class variables”. Grouping by all class variables in R gives exactly that. WithoutNWAY, SAS also outputs the marginal totals (_TYPE_< max), which in R you add with a second summarise over fewer groups oradd_overall()-style helpers.MAXDEC=is display only. Do not round before summarising.- SAS
STDand Rsd()both use the n−1 denominator;VARDEF=changes that in SAS and must be matched if used.
PROC UNIVARIATE and percentiles
proc univariate data=adsl;
class TRT01A;
var AGE;
output out=pct pctlpts=25 50 75 pctlpre=P;
run;
SAS’s default percentile definition is PCTLDEF=5 (empirical distribution function with averaging). R’s quantile() default is type 7 (linear interpolation), and the two differ for small groups.
x <- adsl$AGE[adsl$TRT01A == "Drug"]
rbind(
r_default_type7 = quantile(x, c(.25, .5, .75), type = 7),
sas_pctldef5 = quantile(x, c(.25, .5, .75), type = 2)
) 25% 50% 75%
r_default_type7 54 58 64.75
sas_pctldef5 54 58 65.00
Hyndman and Fan’s type 2 is the SAS default (PCTLDEF=5); type 6 corresponds to PCTLDEF=4 and type 3 to PCTLDEF=2. Put the chosen type in the standards document and in the function signature:
pctl <- function(x, probs = c(.25, .5, .75), sas_def = 5) {
type <- c(`1` = 3, `2` = 3, `3` = 1, `4` = 6, `5` = 2)[as.character(sas_def)]
quantile(x, probs, type = type, na.rm = TRUE, names = FALSE)
}
pctl(x)[1] 54 58 65
Medians agree between type 2 and type 7 for odd n and usually for even n; quartiles are where differences appear. The median of a Q1-Q3 display in a demography table is therefore safe, but the quartiles themselves need the matching type.
Other PROC UNIVARIATE outputs: skewness and kurtosis use the sample-adjusted formulas in SAS (VARDEF=DF); e1071::skewness(type = 2) and kurtosis(type = 2) match them. Normality tests: Shapiro-Wilk is shapiro.test(); SAS reports it only for n ≤ 2000, R up to 5000.
A translation table to keep
| SAS | R | Watch for |
|---|---|---|
PROC FREQ |
count() + mutate(pct = n/sum(n)*100) |
Denominator and MISSING option |
PROC FREQ ... / CHISQ |
chisq.test(correct = FALSE) |
SAS does not apply Yates’ correction by default |
PROC FREQ ... / FISHER |
fisher.test() |
Two-sided p agrees; one-sided definitions differ |
PROC MEANS |
summarise(across(...)) |
n = non-missing; NWAY vs marginals |
PROC UNIVARIATE percentiles |
quantile(type = 2) |
PCTLDEF default 5 |
PROC UNIVARIATE ... NORMAL |
shapiro.test() |
Sample-size limits |
PROC SORT NODUPKEY |
distinct(..., .keep_all = TRUE) |
Keeps first occurrence in both |
PROC TRANSPOSE |
tidyr::pivot_wider()/pivot_longer() |
Variable naming of transposed columns |
Each row of this table is one function in a migration package and one test against a SAS export. That is the whole method: translate once, test once, reuse everywhere. For the full programme see the five-phase migration roadmap and the PROC-to-R equivalence map.