= R Built-in Functions = Several '''functions''' are built into the core R distribution. <> ---- == Mathematical Functions == Most built-in functions operate on [[R/DataTypes#Numeric|numeric]] [[R/Vectors|vectors]] by applying the function element-wise. These include: ||'''Name''' ||'''Meaning''' ||'''Notes'''|| ||`abs` ||absolute value || || ||`all.equal`||Approximate equality of two values||Useful for floating point numbers, but note that it never returns `FALSE`|| ||`ceiling` ||round toward positive infinity || || ||`floor` ||round toward negative infinity || || ||`is.na` ||checks if value is `NA` || || ||`isFALSE` ||checks if value is `FALSE` || || ||`isTRUE` ||checks if value is `TRUE` || || ||`log` ||natural logarithm ||try `log(x, b) for a different base `b`; input of zero gives `-Inf` while input of a negative number gives `NaN`|| ||`log1p` ||natural logarithm of `1 + x` || || ||`log2` ||base-2 logarithm || || ||`log10` ||base-10 logarithm || || ||`round` ||round to nearest integer ||try `round(x,n)` to round to `n` digits|| ||`sqrt` ||square root || || ||`trunc` ||round to 0 || || ---- == Trigonometric Functions == See [[Calculus/Trigonometry#Functions|here]]. ||'''Name'''||'''Meaning''' || ||`sin` ||sine || ||`cos` ||cosine || ||`tan` ||tangent || ||`asin` ||inverse sine || ||`acos` ||inverse cosine || ||`atan` ||inverse tangent|| There is also an `atan2` function; `atan2(y, x)` calculates the angle (in radians) between the x-axis and the vector ''[x y]''. ---- == Statistical Functions == These functions take a (or, at least one) vector that represents observations and return a computed statistic. `NA` values cause most of these functions to also return `NA`. Pass the `na.rm = TRUE` option to ignore these values instead. ||'''Example'''||'''Meaning''' ||'''Notes'''|| ||`cor(a,b)` ||correlation of `a` and `b`|| || ||`cov(a,b)` ||covariance of `a` and `b `|| || ||`density(a)` ||kernel density ||return value is a list|| ||`mad(a)` ||median absolute deviation ||by default, use average of two candidates when number of observations is even; see `low` and `high` options for alternative strategies|| ||`mean(a)` ||mean || || ||`median(a)` ||median || || ||`sd(a)` ||standard deviation || || ||`var(a)` ||variance || || The `sample` function is used as: {{{ > samp.without.replacement <- sample(a, size=5) > samp.with.replacement <- sample(a, size=10, replace=TRUE) > weighted.samp <- sample(a, size=5, prob=b) # b must be same length as a }}} If the `size` option is not passed, then the result is the same size as the input, effectively just performing a randomized sort. Also note that if the input to the `sample` function is a scalar positive integer `n`, then the sample is actually run on the sequence `1:n`. These functions take a (or at least one) matrix where the columns represent vectors of observations. ||'''Name'''||'''Meaning''' || ||`cor(a)` ||correlation matrix of columns in `a`|| ||`cov(a)` ||covariance matrix of columns in `a` || ---- == Distribution Functions == Some of these functions take an output size argument (i.e., `n`). Others take an input that can be a scalar or a vector. The return value matches this input. For the [[Analysis/UniformDistribution|uniform distribution]] from `min` to `max` (defaulting to 0 and 1 respectively): ||'''Example''' ||'''Description''' || ||`runif(n, min = 0, max = 1)`||`n` random values || ||`dunif(x, min = 0, max = 1)`||density at a given `x` value; `1/(max-min)` || ||`punif(q, min = 0, max = 1)`||cumulative probability for a given `q` value || ||`qunif(p, min = 0, max = 1)`||inverse of `punif`; takes a cumulative probability (`p`)|| For the [[Analysis/NormalDistribution|normal distribution]] (defaulting to the standard normal distribution): ||'''Example''' ||'''Description''' || ||`rnorm(n, mean = 0, sd = 1)`||`n` random values || ||`dnorm(x, mean = 0, sd = 1)`||density at a given `x` value || ||`pnorm(q, mean = 0, sd = 1)`||cumulative probability for a given `q` value || ||`qnorm(p, mean = 0, sd = 1)`||inverse of `pnorm`; takes a cumulative probability (`p`)|| For [[Analysis/StudentsTDistribution|Student's t distribution]] with `df` degrees of freedom: ||'''Example'''||'''Description''' || ||`rt(n, df)` ||`n` random values || ||`dt(x, df)` ||density at a given `x` value || ||`pt(q, df)` ||cumulative probability for a given `q` value || ||`qt(p, df)` ||inverse of `pnorm`; takes a cumulative probability (`p`)|| For the [[Analysis/BinomialDistribution|binomial distribution]] with a given `size` and event probability `prob`: ||'''Example''' ||'''Description''' || ||`rbinom(n, size, prob)`||`n` random values || ||`dbinom(x, size, prob)`||density at a given `x` value; the probability of `x` event successes || ||`pbinom(q, size, prob)`||cumulative probability for a given `q` value; the probability at most `q` event successes|| ||`qbinom(p, size, prob)`||inverse of `pbinom`; takes a cumulative probability (`p`) || R does not support functions relating to the [[Analysis/BernoulliDistribution|Bernoulli distribution]]. Instead use the above binomial distribution functions with a size of 1. {{{ rbinom(10, size = 1, prob = 0.7) }}} For the [[Analysis/PoissonDistribution|Poisson distribution]] with a given `lambda` parameter: ||'''Example''' ||'''Description''' || ||`rpois(n, lambda)`||`n` random values || ||`dpois(x, lambda)`||density at a given `x` value || ||`ppois(q, lambda)`||cumulative probability for a given `q` value || ||`qpois(p, lambda` ||inverse of `ppois`; takes a cumulative probability (`p`)|| For the [[Analysis/ChiSquaredDistribution|Chi-squared distribution]] with `df` degrees of freedom: ||'''Example''' ||'''Description''' || ||`rchisq(n, df)`||`n` random values || ||`dchisq(x, df)`||density at a given `x` value || ||`pchisq(q, df)`||cumulative probability for a given `q` value || ||`qchisq(p, df` ||inverse of `ppois`; takes a cumulative probability (`p`)|| Set the random number generator's seed to ensure that random values are reproducible. {{{ set.seed(123) }}} ---- == String Functions == Most built-in functions operate on [[R/DataTypes#Character|character]] [[R/Vectors|vectors]] by applying the function element-wise. These include: ||'''Example''' ||'''Description''' || ||`grepl(pattern,x)` ||Checks if `x` matches || ||`paste(x,y,sep=" ")` ||Concatenate two strings with a separator || ||`paste0(x,y)` ||Concatenate two strings with no separator || ||`strsplit(x,split)` ||Split a string by a regular expression `split` into a list of substrings|| ||`strsplit(x,split,fixed=TRUE)`||Split a string by a literal delimiter `split` || ||`substr(x,start,stop)` ||Extract a substring || ||`toupper(x)` ||Convert to uppercase || ||`tolower(x)` ||Convert to lowercase || Functions that notably do not treat vectors in that manner: * `cat` concatenates and prints the vector * `print` prints the vector * If the `collapse=" "` option is passed on `paste`, then it takes the vectorized output and concatenates into a single string The `sprintf` function provides [[C/StringFormatting|C-style string formatting]]. The format string must be a scalar. The substitution variables can be scalars or vectors, in which case the scalars are recycled as necessary. If multiple substitution variables are vectors, they must be the same length. The `glue` library adds a vectorized function of the same name that supports Python f-string-style interpolation. {{{ library(glue) glue("file_{1:3}.txt") # "file_1.txt" "file_2.txt" "file_3.txt" }}} It also adds a function intended for composing human-readable messages: {{{ glue_collapse(c("foo", "bar", "baz"), sep = ", ", last = ", and ") # "foo, bar, and baz" }}} ---- == Modeling functions == See the [[R/BuiltinFunctions/Lm|lm]] function. A model once fit can be summarized using the `summary` function. The estimated coefficients of a model can be extracted by the `coef` function. The design matrix of a model can be extracted by the `model.matrix` function. A [[Statistics/AnalysisOfVariance|one-way ANOVA]] can be estimated as: {{{ aov(y ~ x, data=df) # If a linear model is already fit: anova(mod) }}} ---- == Plotting functions == See the [[R/BuiltinFunctions/Plot|plot]] function. To create a Q-Q plot of any variable, try: {{{ qqnorm(x, main="Normal Probability Plot of x") }}} To create a histogram from data, use the `hist` function. To create a boxplot from data, try: {{{ boxplot(x) boxplot(x ~ y) # split by levels of y }}} To create a bar chart from statistics, try: {{{ x <- c(3,2,1) names(x) <- c("foo", "bar", "baz") barplot(x) }}} To set two plots side-by-side, try: {{{ par(mfrow=c(1,2)) plot(x,y) plot(x,z) }}} The `par` function can be used this way to create a panel of plots in any arrangement. ---- == Vector Functions == These functions operate on a vector and return a scalar: ||'''Name'''||'''Meaning'''|| ||`length` ||vector length|| These functions operate on a vector and have some side effect, like printing something to the screen, rather than returning a value: ||'''Name'''||'''Meaning''' || ||`head` ||prints the first 6 rows || ||`sort` ||In-place sort a vector || ||`summary` ||prints descriptive statistics|| These functions are useful for working with logical vectors: ||'''Name'''||'''Meaning''' ||'''Notes'''|| ||`all` ||check if every member is `TRUE` ||returns `NA` if any `NA` values present; pass the `na.rm = TRUE` option to ignore the `NA` values|| ||`any` ||check if a member is `TRUE` ||returns `NA` if any `NA` values present; pass the `na.rm = TRUE` option to ignore the `NA` values|| ||`which` ||collects the indices that are `TRUE`|| || To recast other data types, especially lists (which can feature nesting), into vectors, use the `unlist` function. === Equality Testing === Testing for equality of two vectors is tricky. The `all.equal` function does evaluate to `TRUE` in the case that two vectors are equal. If they are not equal however, it evaluates to a text description of the differences, and inconveniently this is a truthy value. Always use the `identical` function to test for equality. === Sequence Generation === The `seq` function has a number of overloaded parameterizations. {{{ seq(start,finish) #equivalent to seq(start,finish,1) seq(finish,start) #equivalent to seq(finish,start,-1) seq(finish) #equivalent to seq(1,finish,1) seq(start, finish, by=0.1) #custom step size seq(start, finish, 0.1) #shorthand, same as above seq(start, finish, length.out=100) #automatically calculates step size }}} The `rep` function also has a number of parameterizations. {{{ rep(1:3, times=2) #1 2 3 1 2 3 rep(1:3, length.out=7) #1 2 3 1 2 3 1 rep(1:3, each=2) #1 1 2 2 3 3 rep(1:3, each=2, length.out=7) #1 1 2 2 3 3 1 }}} === Sorting === The `sort` function has a few options. * To sort descending instead of ascending, pass `decreasing = TRUE`. * To retain `NA` values and sort them to the end, pass `na.last = TRUE`. Or to retain them and sort them to the beginning, set as `FALSE`. The default is for `NA` values to be removed. ---- == Matrix functions == The `replicate` function collects the results of a repeated command into a matrix. For example, to generate 10 samples of 100 observations each from the standard normal distribution, with each sample becoming a column, try: {{{ replicate(10, rnorm(100)) }}} The `apply` function applies a function to a rows and/or columns of a matrix. {{{ row_sums <- apply(x, MARGIN=1, FUN=sum) col_sums <- apply(x, MARGIN=2, FUN=sum) all_sums <- apply(x, MARGIN=c(1,2), FUN=sum) }}} ---- == See also == [[https://www.rdocumentation.org/packages/base|R documentation on functions in the base package]] [[https://www.rdocumentation.org/packages/stats|R documentation on functions in the stats package]] [[https://www.rdocumentation.org/packages/graphics|R documentation on functions in the graphics package]] ---- CategoryRicottone