Dual Programming SAS vs R: Numerical Tolerance Rules That Survive QC
clinical
sas-to-r
validation
How to compare SAS and R outputs at dataset and TLF level during a migration: which tolerances to use for p-values, estimates and percentages, how to automate the comparison in R, and how to classify every difference as explained, accepted or deviation.
Author
Rverse Analytics
Published
October 8, 2026
Compare at three levels: dataset (structure, keys, values), TLF (every displayed cell) and log (warnings, record counts). Each level has its own tolerance.
Use absolute tolerance for bounded quantities (percentages, p-values, correlations) and relative tolerance for unbounded estimates (means, hazard ratios, confidence limits). Typical values: 1e-8 for values stored in full precision, half a unit in the last displayed digit for displayed values.
Every difference gets one of three labels: explained (convention, e.g. rounding), accepted (within tolerance) or deviation (investigate). The counts of each go into the validation summary.
all.equal() and waldo::compare() cover most dataset comparisons; a small wrapper gives you a PROC COMPARE-style report.
Dual programming is the backbone of a SAS-to-R migration: one programmer produces the output in R, the SAS original stands as the reference, and the two are compared programmatically. The comparison is only useful if the rules for “same” are written down before anyone looks at the numbers. This post is the set of rules we start from.
Three comparison levels
1. Dataset level. Analysis datasets (ADaM or sponsor-specific) must match in variable names, types, labels where required, sort keys, record counts and values. Values are compared variable by variable with a tolerance appropriate to the type.
2. TLF level. Tables are compared cell by cell on the displayed string after both sides apply the same rounding rule. Figures are compared on the data behind them (the plotted coordinates), not pixel by pixel.
3. Log level. Record counts at each derivation step, number of subjects per treatment arm, and the absence of unexpected warnings. This catches silent filtering differences that value comparison alone can miss.
The weaker tolerances for iterative models are not a weakness of R. SAS PROC MIXED and R nlme/lme4 use different optimisers and stopping rules; agreement to four significant digits is what the mathematics gives you unless you tighten both sides.
Automating the dataset comparison
all.equal() already handles numeric tolerance. A thin wrapper turns it into a report you can file:
compare_ds <-function(base, compare, keys, tol =1e-8) {stopifnot(all(keys %in%names(base)), all(keys %in%names(compare))) base <- base[do.call(order, base[keys]), , drop =FALSE] compare <- compare[do.call(order, compare[keys]), , drop =FALSE] vars <-intersect(names(base), names(compare)) out <-lapply(vars, function(v) { b <- base[[v]]; c <- compare[[v]]if (is.numeric(b) &&is.numeric(c)) { diff <-abs(b - c); diff[is.na(b) &is.na(c)] <-0data.frame(variable = v, type ="numeric", n_diff =sum(diff > tol, na.rm =TRUE),max_abs_diff =suppressWarnings(max(diff, na.rm =TRUE))) } else { same <- (as.character(b) ==as.character(c)) | (is.na(b) &is.na(c))data.frame(variable = v, type ="character", n_diff =sum(!same, na.rm =TRUE), max_abs_diff =NA) } }) res <-do.call(rbind, out)attr(res, "only_in_base") <-setdiff(names(base), names(compare))attr(res, "only_in_compare") <-setdiff(names(compare), names(base))attr(res, "n_records") <-c(base =nrow(base), compare =nrow(compare)) res}# Simulated SAS reference vs R output for an ADSL-like datasetset.seed(7)sas <-data.frame(USUBJID =sprintf("01-%03d", 1:8), TRT01A =rep(c("Placebo", "Drug"), 4),AGE =c(54, 61, 47, 70, 66, 58, 49, 63), BMIBL =round(rnorm(8, 27, 3), 1))r <- sasr$BMIBL[3] <- r$BMIBL[3] +1e-10# floating-point noise: acceptedr$AGE[5] <-65# a real deviationres <-compare_ds(sas, r, keys ="USUBJID")res
variable type n_diff max_abs_diff
1 USUBJID character 0 NA
2 TRT01A character 0 NA
3 AGE numeric 1 1.000000e+00
4 BMIBL numeric 0 9.999823e-11
attr(res, "n_records")
base compare
8 8
For interactive work, waldo::compare(sas, r, tolerance = 1e-8) gives a coloured diff of exactly the same information, and diffdf::diffdf() produces a PROC COMPARE-like listing if you prefer a dedicated package.
Comparing TLFs
Treat each table as a long data frame: one row per (table, row label, column label, cell string). Produce that structure from both the SAS RTF/listing and the R output (we parse the SAS RTF or, better, ask for the SAS summary dataset behind the table). Then:
tlf_sas <-data.frame(row =c("Age, mean (SD)", "Female, n (%)", "BMI, mean (SD)"),col ="Drug", cell =c("61.3 (8.2)", "3 (37.5)", "27.4 (2.9)"))tlf_r <-data.frame(row = tlf_sas$row, col ="Drug",cell =c("61.3 (8.2)", "3 (37.5)", "27.4 (3.0)"))m <-merge(tlf_sas, tlf_r, by =c("row", "col"), suffixes =c("_sas", "_r"))m$match <- m$cell_sas == m$cell_rm
row col cell_sas cell_r match
1 Age, mean (SD) Drug 61.3 (8.2) 61.3 (8.2) TRUE
2 BMI, mean (SD) Drug 27.4 (2.9) 27.4 (3.0) FALSE
3 Female, n (%) Drug 3 (37.5) 3 (37.5) TRUE
A mismatch in a displayed SD at the last digit is the classic case for the three-way classification below.
Classifying differences
Every non-match receives one label and one sentence:
Explained: a documented convention difference. Example: “SD displayed 2.9 vs 3.0; SAS rounds half away from zero, R rounds half to even; stored values agree to 1e-10.” These are closed by reference to the tolerance rules.
Accepted: within the stated tolerance but not identical. Example: hazard ratio 0.7312 vs 0.7311 (relative difference 1.4e-4 under a 1e-4 rule for mixed models would not be accepted; under a Cox model rule of 1e-6 it is a deviation). The example shows why rules must be tied to the method.
Deviation: outside tolerance or not explained. Each deviation has an owner, a root cause and a resolution (fix R, fix SAS reference, or amend the specification). Deviations are never closed by widening the tolerance after the fact.
The validation summary reports the count of comparisons, the count in each class, and a listing of deviations with their resolutions. A table package with, say, 4,800 cells compared, 4,788 identical, 11 explained and 1 resolved deviation is a strong, honest result.
Common sources of deviations we see
Population flags derived differently (e.g. ITT vs safety) because a SAS macro applied a filter that was not in the specification.
Tie handling in Wilcoxon and log-rank tests; survival::survdiff() and coin::wilcox_test() can be configured to match SAS defaults.
Type III vs Type I sums of squares in ANOVA. R’s anova() on lm() is sequential (Type I); use car::Anova(type = 3) with appropriate contrasts.
Missing-value sorting and comparisons, where SAS treats missing as the smallest value. See missing values in SAS vs R.
Date arithmetic, especially partial dates imputed by a macro.
Each of these becomes a line in the equivalence report the first time it appears and a rule in the standards the second time.