Skip to main content

From SAS Macros to R Functions: A Worked Conversion

clinical
sas-to-r
r-programming
A SAS summary macro converted step by step into a tested R function inside a package: parameter mapping, by-group handling, missing values, formatted output and unit tests. The pattern that makes a SAS-to-R migration maintainable.
Author

Rverse Analytics

Published

October 8, 2026

  • A SAS macro is text substitution; an R function is a first-class object. Convert the intent (inputs, outputs, edge cases), not the tokens.
  • Map macro parameters to function arguments with defaults, replace %IF branches with vectorised logic, and return a data frame rather than writing a dataset to a library.
  • Put converted functions in an internal package with roxygen documentation and testthat tests. Each test is a reusable piece of validation evidence.
  • One shared function replaces N copies of a macro. That is where the migration’s cost savings come from.

Shared macros are the assets in a SAS codebase. They are also where SAS-to-R migrations go wrong when programmers translate %DO loops into for loops and CALL SYMPUT into global variables. Here is a small but realistic conversion done the way we do it in a migration package.

The SAS macro

A typical descriptive-statistics macro: summarise a numeric variable by treatment, returning n, mean (SD), median and range, formatted for a table.

%macro desc_stat(indata=, var=, by=TRT01A, dec=1, out=);
  proc means data=&indata noprint nway;
    class &by;
    var &var;
    output out=_s n=n mean=mean std=sd median=med min=min max=max;
  run;
  data &out;
    set _s;
    length stat $40;
    n_c    = put(n, 3.);
    meansd = cats(put(mean, 8.&dec), " (", put(sd, 8.%eval(&dec+1)), ")");
    med_c  = put(med, 8.&dec);
    range  = cats(put(min, 8.&dec), ", ", put(max, 8.&dec));
    keep &by n_c meansd med_c range;
  run;
%mend desc_stat;

%desc_stat(indata=adsl, var=AGE, out=t_age);

Things to notice: the by-variable is a parameter; decimal places drive formats for mean (dec) and SD (dec + 1); the output is a dataset of character columns ready for the table layer; missing values are silently dropped by PROC MEANS.

The R function

library(dplyr)

round_sas <- function(x, digits = 0) sign(x) * floor(abs(x) * 10^digits + 0.5 + 1e-9) / 10^digits

fmt <- function(x, dec) {
  ifelse(is.na(x), "", formatC(round_sas(x, dec), format = "f", digits = dec))
}

desc_stat <- function(data, var, by = "TRT01A", dec = 1, na_rm = TRUE) {
  stopifnot(is.data.frame(data), var %in% names(data), by %in% names(data))
  data |>
    group_by(across(all_of(by))) |>
    summarise(
      n      = sum(!is.na(.data[[var]])),
      mean   = mean(.data[[var]],   na.rm = na_rm),
      sd     = sd(.data[[var]],     na.rm = na_rm),
      med    = median(.data[[var]], na.rm = na_rm),
      min    = suppressWarnings(min(.data[[var]], na.rm = na_rm)),
      max    = suppressWarnings(max(.data[[var]], na.rm = na_rm)),
      .groups = "drop"
    ) |>
    transmute(
      across(all_of(by)),
      n_c    = formatC(n, width = 3),
      meansd = paste0(fmt(mean, dec), " (", fmt(sd, dec + 1), ")"),
      med_c  = fmt(med, dec),
      range  = paste0(fmt(min, dec), ", ", fmt(max, dec))
    )
}

set.seed(11)
adsl <- data.frame(
  USUBJID = sprintf("01-%03d", 1:40),
  TRT01A  = rep(c("Placebo", "Drug 10 mg"), each = 20),
  AGE     = c(round(rnorm(20, 58, 9)), round(rnorm(20, 60, 8)))
)
adsl$AGE[c(3, 27)] <- NA
desc_stat(adsl, "AGE")
# A tibble: 2 × 5
  TRT01A     n_c   meansd      med_c range     
  <chr>      <chr> <chr>       <chr> <chr>     
1 Drug 10 mg " 19" 57.6 (5.90) 57.0  47.0, 72.0
2 Placebo    " 19" 55.6 (7.42) 55.0  44.0, 70.0

What changed, and why

  • Parameters became arguments with defaults. by = "TRT01A" and dec = 1 mirror the macro, and stopifnot() gives an immediate, readable error if a column is missing, where the macro would have produced a cryptic log note.
  • Formats became a helper. fmt() centralises rounding and width. We use SAS-style rounding here because the function is for the parallel-run phase; see rounding differences.
  • Missing values are explicit. n counts non-missing, matching PROC MEANS, and the na_rm argument makes the policy visible instead of implicit.
  • No side effects. The function returns a data frame. Writing to disk is the caller’s decision, which is what makes the function testable.

Putting it in a package

In the migration package (usethis::create_package("cromigr")), the function gets roxygen documentation and a test file:

# tests/testthat/test-desc_stat.R
test_that("desc_stat matches SAS reference for AGE", {
  ref <- readRDS(test_path("fixtures", "t_age_sas.rds"))   # exported from SAS
  out <- desc_stat(adsl_fixture, "AGE")
  expect_equal(out, ref)
})

test_that("desc_stat handles all-missing groups without error", {
  d <- data.frame(TRT01A = c("A", "B"), AGE = c(NA_real_, NA_real_))
  expect_no_error(desc_stat(d, "AGE"))
})

The first test is validation evidence: it encodes the SAS reference output and proves the R function reproduces it. The second protects against the edge case that would have surfaced as a warning in a SAS log that nobody read. devtools::test() runs them on every change; renv pins the package set.

Patterns for other macro types

SAS macro pattern R equivalent
%DO i = 1 %TO &n over variables lapply()/purrr::map() over a character vector of names
%IF &type = CAT %THEN branches Separate small functions, dispatched by an argument or S3 class
CALL SYMPUT to pass values around Return values; never global assignment
PROC SQL joins inside macros dplyr::left_join() with explicit by =
Macro writing many datasets to WORK A list of data frames, or one long data frame
%INCLUDE of shared code library(cromigr)

The general rule: anything the macro does by rewriting text, the function should do with data structures.

What this buys you

Fifty studies that each called %desc_stat now call one tested function. A specification change is one edit, one test update and one package version. The package’s test log and riskmetric scores feed the validation plan directly. This is Phase 2 of the five-phase roadmap, and it is where the migration either pays for itself or does not.

Related: PROC FREQ, PROC MEANS and PROC UNIVARIATE in R, and the SAS to R migration service.