Browse all practice questions for the Casualty Actuarial Society MAS-1 Practice Exam. Search by topic, open any question and review its full explanation, then test yourself in the practice quiz.

Casualty Actuarial Society MAS-1 Practice Exam 2026 - Free MAS-1 Practice Questions and Study Guide course image
More practice questions

These questions are part of the practice quiz. Start practicing

  • Given Y claims are made by time t, the unordered times of past events follow what distributions?
  • Which statement about a simple linear relationship is true?
  • Ridge regression coefficients are not scale equivariant.
  • What is the canonical link for the Bernoulli distribution in GLMs?
  • Which norm is used in the penalty term for lasso regression?
  • If we want the expected number of time periods a chain is spent in state j given it started in state i, what are the steps for calculating this?
  • PCA can be used for data visualization.
  • In ridge regression, the sum of squares of beta is bounded above by s. Which statement best captures this constraint?
  • What is the canonical link function for Poisson regression?
  • Performing k-fold cross validation requires fitting a model for a total of k times.
  • The Cramer-Rao lower bound for the variance of all unbiased estimators of theta equals:
  • For X ~ Exp(λx) and Y ~ Exp(λy) with 1 < X < Y, what is E[X | 1 < X < Y]?
  • R^2 is the fraction of variation in y about the mean of y that's explained by the linear relationship with x.
  • Which expression represents Cov(X,Y) in terms of expectations?
  • Which distribution is associated with the canonical link that is inverse?
  • Which of the following best describes KNN?
  • True or false: since training error can be a poor estimate of the test error, RSS and R-squared are not suitable for selecting the best model.
  • What is the canonical link for the gamma/exponential distribution?
  • SSE equal to 0 indicates overfitting.
  • what is the excel equation for solving for the pdf in gaussian kernel estimation?
  • For a sample from an inverse Gaussian distribution, which expression is the MVUE of the mean parameter?
  • In an exponential family distribution, the sufficient statistic for the natural parameter mu, given observations x1,...,xn, is which of the following?
  • In Gaussian kernel density estimation, the data point x_i represents what in the kernel sum?
  • The fewer positive raw moments that exist, the greater the tail weight.
  • Which statement best indicates a variable is statistically significant in a standard hypothesis test?
  • When forming a confidence interval for a proportion, which statistic is used?
  • For an exponential distribution with complete data, what is the maximum likelihood estimator of the mean parameter?
  • In a fitted values vs residuals graph, heteroscedasticity is indicated by which pattern?
  • what is the neyman-pearson theorem for hypothesis testing?
  • In a PCA performed on a data set with 50 observations and 3 independent continuous variables, which statement is true?
  • The Excel function CHISQ.DIST.RT can be used to compute the p-value for a chi-square statistic.
  • Removing any component from a minimal path set will break the guarantee.
  • What is the test statistic for testing the equality of two variances?
  • When forming a confidence interval for the difference between two means with paired observations, which statistic is used?
  • The Cramer-Rao lower bound for the variance of unbiased estimators is given by which expression?
  • Removing high leverage points in a linear regression model primarily affects which aspect of the fitted model?
  • Collinearity reduces the accuracy of the estimates of regression coefficients and may make it harder to reject the hypothesis that beta_j = 0.
  • In dummy coding for a categorical predictor with four levels, how many dummy variables are needed if one category is used as the baseline?
  • If X ~ Uniform(m, n), what is the distribution of (X | X > pi_q)?
  • Which statement correctly describes the relationship between MSE and the true parameter?
  • Forward stepwise selection cannot be used in high-dimensional settings.
  • Poisson regression models assume exposure is constant when modeling the rate at which events occur.
  • ANOVA is a useful approach for analyzing the means of groups of continuous response variables, where the groups are categorical.
  • What is the MVUE of a normal distribution with a defined variance?
  • What does removing outliers do to a linear regression model?
  • GAMs are a useful representation if we are interested in inference, since you can examine the effect of the predictor variables on the response while holding all of the other predictor variables constant.
  • Are all Poisson processes characterized by stationary and independent increments?
  • What is the MVUE of a Poisson distribution?
  • In a residuals vs fitted values plot, heteroscedasticity is indicated by which pattern?
  • What is the expected value of the top q% of losses?
  • After standardizing the predictors, which statement about the PLS first direction is correct?
  • In the context of Poisson residuals, the Pearson residual is defined using which denominator?
  • Which distribution has a constant hazard rate?
  • What is the canonical link for the Poisson distribution?
  • PCR assumes that the directions in which features show the most variation are the directions that are associated with the target.
  • K-fold validation has an advantage over LOOCV in variance reduction.
  • Deviance is a useful measure of goodness of fit for all models in the exponential family.
  • For a given dataset, the number of variables in a Lasso regression model will always be greater than or equal to the number of variables in a Ridge regression model.
  • A logit model applies when the explanatory variables are both continuous and categorical.
  • What represents the moment generating function of Y evaluated at t=1, My(1)?
  • With a cubic spline, what must match at the knot?
  • What makes a good argument for choosing LOOCV over 5-fold CV?
  • For paired observations, which statistic is used to form a confidence interval for the difference in means?
  • Which Poisson process type has stationary increments?
  • An unbiased estimator is considered a consistent estimator if the variance of the estimator converges to 0 as n approaches infinity.
  • If both classes were transient, after some time, the chain would not be in either class.
  • The cumulative proportion of variance explained cannot decrease when more PCs are added.
  • Which statement about training set MSE versus test MSE is true?
  • Under X|θ ~ N(θ, σ^2) and θ ~ N(μ, τ^2), the marginal distribution of X is Normal with mean μ and variance σ^2 + τ^2.
  • In a one-dimensional symmetric random walk, where the probability of moving in either direction is 0.5, are all states recurrent?
  • True or false: Local regression is a memory-based procedure.
  • The pseudo R-squared is computed as one minus the ratio of the model log-likelihood to the null log-likelihood.
  • LOOCV uses n training/testing splits, one for each observation left out.
  • For modeling hourly bike-sharing usage by day of week, which distribution and link function are most appropriate?
  • In Poisson regression, the Pearson residual is computed as (y - mu_hat) / sqrt(mu_hat). What does mu_hat represent in this context?
  • True or false: A large value of Mallows Cp indicates a model with a high test error.
  • Which statement describes a common effect of high dimensionality on MSE estimates?
  • In a Markov chain, positive recurrence is a property that holds for all states in a given communicating class.
  • LOOCV is a special case of k-fold cross-validation.
  • Power parameter in the Tweedie family that corresponds to a Poisson distribution.
  • Which type of qualitative variable has categories with a meaningful order?
  • In PCA, the third principal component is orthogonal to the first principal component.
  • In supervised learning, the variance and the squared bias are inversely related.
  • The deviance is defined as a measure of distance between saturated and fitted model.
  • In a Markov chain, a state whose probability of returning is less than 1 is called a
  • For a sample from Uniform(a,b), the expected minimum (k = 1) is a + (b - a)/(n + 1). Which expression correctly represents this?
  • What is the sum of the leverages across all observations in a linear regression with an intercept and p predictors?
  • Unlike the validation set approach, the k-fold cross-validation approach uses all observations to train the model.
  • Which model form will both have a discontinuity in the fitted curve and most likely overfit the data when predicting height from shoe size?
  • Which of the following statements best describes a non-homogeneous Poisson process?
  • Among the following kernel options for kernel density estimation, which are symmetric?
  • True or False: R-squared is a good measure for model comparison.
  • Which statement best describes how lasso influences model sparsity?
  • Is the density f(y; θ) = θ y for y > θ a member of the exponential family?
  • In PCA, the first principal component is the direction along which the data vary the most.
  • k-fold cross-validation has higher variance than LOOCV when k<n.
  • Is a regression tree an example of supervised or unsupervised learning?
  • PLS is a subset selection method.
  • Which Excel expression yields the p-value for a chi-square test?
  • Can a deviance plot be used to visually approximate the MLE of theta?
  • K-fold validation has an advantage over LOOCV in bias reduction.
  • In a regression with p predictors and an intercept, the trace of the hat matrix (sum of leverages) equals:
  • Which statement about the relationship between Mallows Cp and AIC is most consistent with the material?
  • If the hazard rate function decreases with x, the distribution has a heavy tail.
  • SSE equals zero indicates overfitting.
  • Is KNN an example of supervised or unsupervised learning?
  • In ridge regression, which parameter is not subject to shrinkage?
  • The squared bias increases as the method's flexibility decreases.
  • Homoscedasticity occurs when what condition holds?
  • Which statement best captures the meaning of stationary and independent increments?
  • The deviance for normal distributions is proportional to the residual sum of squares.
  • In ridge regression, coefficients shrink toward zero but are typically not exactly zero.
  • In ridge regression, irreducible error as s increases?
  • When a lognormal X ~ Lognormal(mu, sigma^2) is scaled by a positive constant c, which distribution describes cX?
  • Ordinary least squares estimators are inherently unbiased.
  • When forming a confidence interval for the difference of two proportions, which statistic is used?
  • Which statement about lasso regression compared to ordinary least squares is true?
  • Mallows Cp is an unbiased estimate of the test MSE if calculated using an unbiased estimate of variance.
  • Loadings for the first direction are proportional to covariances between the response and each standardized predictor.
  • What happens to the variance-covariance matrix when implementing the quasi-likelihood method?
  • In generalized linear models, the statement 'the saturated model has the highest possible deviance' is true or false?
  • Which of the following statements is true about one-dimensional and two-dimensional symmetric random walks?
  • Cross-validation is used to measure the accuracy of a parameter estimate.
  • Lasso regression performs variable selection by shrinking some coefficients exactly to zero.
  • What is MSE(estimator)?
  • In a normal linear model, the scaled deviance is equal to which of the following?
  • How do you test for time reversibility of a Markov chain?
  • Which term describes a Markov chain that has only one communicating class?
  • In the transient-state fundamental matrix S = (I - PT)^{-1}, what does the entry S_ij represent?
  • A Poisson process with a constant rate is called what?
  • In the kernel density estimator context, the contribution of a single kernel centered at x_i to the pdf is proportional to which expression?
  • How do you identify overdispersion in a model?
  • If an explanatory variable is uncorrelated with all other explanatory variables, the corresponding VIF would equal what?
  • Greedy Algorithm B uses which sequence of k values?
  • In local regression, increasing the span parameter makes the fit more global rather than local.
  • The MVUE is defined as the unbiased estimator with the minimum variance.
  • For a Gamma distribution with known shape parameter alpha and i.i.d. samples, what is the MVUE of the scale parameter beta?
  • If the number of claims follows a Poisson distribution with rate lambda, what distribution describes the waiting time until the first claim?
  • What is the MVUE of sigma^2 for a normal distribution with unknown mean and unknown variance?
  • Which interval quantifies the possible range for a future observation Y given X?
  • What is the symbol used for the expected number of time periods a chain is in state 2 given the chain starts in state 1?
  • In kernel density estimation, increasing the bandwidth reduces variance but can increase bias; the statement about smoother pdf holds true when bandwidth is larger.
  • What is the MVUE of a binomial distribution?
  • Under a saturated model, the predicted value for a given observation is
  • The first PC is the line in p-dimensional space that is closest to the observations.
  • In a series system, the system functions only if every component is functioning.
  • In computing the first direction, PLS places the highest weight on the variables that are most strongly related to the response.
  • Which method is used to select the appropriate level of model flexibility?
  • In least squares regression, the LOOCV estimate for the test MSE can be calculated by fitting a model once.
  • Which of the following is generally considered supervised learning?
  • Ridge regression uses an L2 penalty, while lasso regression uses an L1 penalty.
  • In kernel density estimation, what does the cdf indicate we are looking for?
  • Which expression is the MSE of an estimator?
  • Residual sum of squares is a suitable metric for selecting the best model among models with different numbers of predictors.
  • A series system functions only when all components function.
  • For a negative binomial distribution with parameter r, the MVUE of the scale parameter is:
  • PCR is useful for performing feature selection.
  • In a scenario where k-out-of-n with k=1 and n=5, the expected lifetime equals theta times the harmonic sum 1 + 1/2 + 1/3 + 1/4 + 1/5. If theta=2, which value is correct?
  • Adjusted R-squared equal to 1 indicates overfitting.
  • Which statement about LOESS span and smoothing is true?
  • Which step is typically performed first when computing expected sojourn times for a Markov chain with transient states?
  • For X ~ Uniform(2,5), conditioning on X > 3 yields X | X > 3 ~ Uniform(3,5).
  • What is the formula for a Pearson residual?
  • LOOCV requires fitting a model a total of n times.
  • In kernel density estimation, after calculating the kernel contributions k_i(x) from each observation, how is the density estimate at x formed?
  • How can you identify the mode using a probability density function?
  • Which description matches the constrained form used in the lasso alternative objective?
  • If there is a correlation among the error terms, then the estimated standard errors will tend to underestimate the true standard errors.
  • Both lasso and ridge regression shrink coefficients toward zero, but lasso can force some coefficients to be exactly zero.
  • The Tweedie distribution is particularly useful when data include zeros and continuous positive values that can be viewed as which distribution?
  • LOOCV corresponds to k-fold cross-validation with k equal to n.
  • In simple linear regression, R^2 equals the square of the sample correlation coefficient r. Which option is true?
  • In ridge regression, the shrinkage penalty is applied to all coefficient estimates except for the intercept.
  • What is the test statistic for a likelihood ratio test?
  • Which sequence of steps correctly calculates the probability of moving from transient state 2 to state 4 in a Markov chain?
  • If N ~ Poisson(lambda) and S = sum_{i=1}^N X_i where X_i are independent of N with E[X^2] finite, what is Var(S)?
  • What is the only continuous distribution listed with a finite support?
  • Does the quasi-likelihood method change the coefficient estimates?
  • R^2 equal to 1 indicates overfitting.
  • An annuity-due issued to a 40-year-old that pays 1 each year until death or age 60, whichever comes first, has an actuarial present value given by which expression?
  • For a normal distribution, the deviance is proportional to the residual sum of squares.
  • When scaling X ~ Lognormal(mu, sigma^2) by a positive constant c, which distribution describes cX?
  • The irreducible error variance in this model is 100^2.
  • Why do we use w-1 dummy variables for a categorical predictor with w levels?
  • Explanatory variables for Poisson regression besides exposure can be either continuous or categorical.
  • Which distribution and link function should be used for slices of pizza sold at a convenience store based on distance to the city center?
  • In Tweedie distributions, data are appropriate when data include zeros and continuous positive values that can be viewed as which distribution?
  • How is Greedy Algorithm A described for optimization problems?
  • Shrinkage fits a model involving a subset of predictors with the estimated coefficients shrunken towards zero.
  • Which statement about Mallows Cp and AIC is supported by the material?
  • Lasso regression tends to outperform ridge regression in terms of bias, variance, and MSE.
  • If two sampling distributions have the same mean, the one with smaller variance is called what?
  • Ridge regression shrinks the coefficient estimates, which has the benefit of reducing the bias.
  • Which ordering of the degrees of freedom used by the three spline models is correct from most to least, given a linear spline with k knots, a cubic spline with k knots, and a natural cubic spline with k total knots (k minus 2 interior knots)?
  • For a distribution in the exponential family, E[Y] equals negative derivative ratio - c'(θ) / b'(θ).
  • When forming a confidence interval for a mean with unknown variance, which statistic is used in the calculation?
  • Given the minimal path sets {1,2,5}, {1,3,4}, {2,3,5}, {3,4,5}, which of the following is a minimal cut set?
  • In the Tweedie family, p = 0 corresponds to which distribution?
  • In the exponential family, the expression for E[Y] can be written as E[Y] = - c'(θ) / b'(θ).
  • What is Mallows Cp equation?
  • In simple and multiple linear regression, the maximum likelihood estimator for the regression coefficients coincides with the ordinary least squares estimates when residuals are normally distributed.
  • When should you use pooled variances for confidence intervals (or tests)?
  • Which formula correctly expresses Var(X+Y) accounting for dependence?
  • Can a plot of the score function be used to visually approximate the MLE of theta?
  • Ridge regression shrinks coefficients toward zero and can never set any coefficient exactly to zero.
  • In simple linear regression, the statement that the sample correlation between x and y equals the coefficient of determination R^2 is
  • The number of events that occur in disjoint time intervals must be independent is a property of counting processes.
  • In Lasso regression, as lambda increases, what happens to the variance of the predictions?
  • Which statement about the exponential family and canonical form is true?
  • Which method is used when a problem mentions the 'best critical region'?
  • In ridge regression, training error as s increases?
  • Using an alternative fitting procedure makes results easier to interpret.
  • Using an alternative fitting procedure will likely result in a simpler model.
  • In marketing analytics, clustering to segment shoppers is best described as which type of learning?
  • If a distribution isn't in canonical form, can there be a natural parameter?
  • Which of the following best describes the objective of lasso regression?
  • In a Poisson process, increments over disjoint time intervals are what property?
  • Subset selection is used to identify a subset of the predictors and then fit a model using least squares on the reduced set of variables.
  • What is the probability that an observation is not selected for a bootstrap sample?
  • If state 2 is positive recurrent, then state 4 must be positive recurrent.
  • Deviance is minimized to obtain the best-fitting generalized linear model; in general, lower deviance indicates a better fit.
  • Leave-one-out cross-validation is a special case of k-fold cross-validation where k equals the number of observations.
  • Which methods guarantee a nested sequence of models as predictors are added or removed?
  • In regression with Gaussian errors, a large value of Mallows' Cp indicates a model with a low test error.
  • LOOCV bias is lower than k-fold CV bias.
  • Cluster analysis is typically categorized as which type of learning?
  • Mean squared error (MSE) is defined as the expected squared difference between the estimator and the true parameter. Which statement is true?
  • Which of the following best describes a typical effect of the L1 penalty in lasso regression?
  • The standard error of regression uses degrees of freedom equal to n-2.
  • Dimension reduction involves projecting into a lower-dimensional subspace using M linear combinations.
  • To model a non-negative response with an unbiased estimate, which error structure and link function combination is most appropriate?
  • In linear regression with p predictors and an intercept, the sum of leverages equals which of the following?
  • Which of the following is generally considered unsupervised learning?
  • The smoothness of a continuous predictor variable in a GAM can be summarized by degrees of freedom.
  • Which of the following increases monotonically as model flexibility increases?
  • How is overdispersion detected in a generalized linear model?
  • In simple linear regression, a random pattern in the scatterplot of y against x indicates that R^2 is near zero.
  • Only w-1 dummy variables are needed to represent w classes of a categorical predictor.
  • For X ~ Uniform(m,n), the conditional distribution X | X > pi_q is Uniform(pi_q, n) provided pi_q lies in (m,n).
  • Adjusted R^2 equal to 1 indicates overfitting.
  • Backward stepwise selection cannot be performed on the dataset if n<p.
  • If X_(k) is the kth order statistic from an iid sample from Uniform(0, θ), which distribution does X_(k) follow?
  • True or false: Mallows Cp is an unbiased estimate of the test MSE if its variance is calculated using an unbiased estimate of the variance.
  • A probit link is a valid alternative to the logistic link for binary outcomes.
  • The variance of error terms doesn't have to be constant.
  • The training MSE decreases as model flexibility increases.
  • What is the form of the likelihood function for two independent populations with different density parameters?
  • A consistent estimator is also unbiased.
  • The law of total variance states Var(X) = E[Var(X|θ)] + Var[E(X|θ)].
  • Poisson regression assumes that the mean equals the variance of the response variable.
  • Poisson regression models incorporate a logarithmic link function.
  • Which statement about k-fold cross-validation is true?
  • A Markov chain that has a limiting distribution is described as which type?
  • In ordinary least squares, the variance of each coefficient estimate is given by which expression?
  • In ridge regression, test error as s increases?
  • Which expression expresses E[Y] for a distribution in the exponential family?
  • A GLM with the same distribution and link function as the model of interest that has the max number of parameters that can be estimated; assists in assessing model adequacy is called what?
  • Using Cook's distance with a unity threshold, an observation is influential if which condition holds?
  • In ridge regression, increasing the tuning parameter lambda strengthens the penalty and shrinks coefficients toward zero.
  • What is the canonical link for the normal distribution?
  • A chain with only one class is called an irreducible chain.
  • What is MGF of Y evaluated at t=1?
  • Residual plots are a useful graphical tool for identifying non-linearity.
  • If n = 4, how many minimal path sets are there?
  • Best subset selection requires fitting all (2 choose p) models for each possible combination of p predictors.
  • In life-table notation, p_x denotes the probability of surviving from age x to age x+1.
  • Increasing the significance level would increase the probability of a Type I error.
  • In the Poisson distribution, which statement is true about the relationship between the mean and the variance?
  • In a Markov chain, a state that cannot be left once entered is called an absorbing state.
  • Which of the listed modeling procedures performs variable selection?
  • In a smoothing spline model fit to data, what happens to bias as the tuning parameter lambda increases?
  • When deciding between a regression spline and local regression, which component must be considered for a regression spline but not for local regression?
  • In general, LOOCV requires fitting a model for a total of n times.
  • In GLMs, the primary consideration for choosing between a Poisson model with a log link and a Gaussian model with an identity link is the distribution of the response variable.
  • If two statistics have the same mean, the one with smaller variance is called the efficient estimator.
  • True or false: Ridge regression is less flexible and thus results in an improved prediction accuracy when its decrease in variance is less than its increase in squared bias.
  • Evaluating the correlation matrix of predictor variables is a reliable method to detect collinearity.
  • Lasso regression is able to perform variable selection by forcing some coefficients to be exactly zero.
  • Rank the following tools by flexibility in descending order: spline, linear regression, ridge regression.
  • What is the expected value of the kth order statistic from a Uniform(a,b) distribution?
  • In Greedy Algorithm A, after selecting the lowest-cost assignment, what is done next?
  • For an exponential random variable X, what is E[X | X > a]?
  • Which statement about lasso regression is false?
  • In a linear model with an intercept and p explanatory variables, the leverage for each observation must be between which values?
  • Which statement correctly differentiates homogeneous and non-homogeneous Poisson processes regarding stationary increments?
  • Which statement best describes the difference between homogeneous and non-homogeneous Poisson processes?
  • A saturated model has a deviance of zero.
  • Using an alternative fitting procedure will likely improve prediction accuracy.
  • If n = 4, how many minimal cut sets are there?
  • In a three-state Markov chain with states 0, 1, 2 and starting in 0, what is the formula for the expected number of steps to return to state 0?
  • Consistency of an estimator is characterized by the variance converging to zero as the sample size grows.
  • Which type of qualitative variable has categories without a meaningful order?
  • Decreasing the significance level would increase the probability of a Type II error.
  • In simple linear regression, the least squares line passes through the point (x-bar, y-bar).
  • Which distribution uses the canonical link inverse squared?
  • In a Galton-Watson branching process with offspring probabilities P_j, the extinction probability π0 satisfies which equation?
  • Power is defined as 1 minus the probability of a Type II error.
  • Bootstrapping can be used to select the appropriate level of model flexibility.
  • Dimension reduction involves projecting the p predictors into an M-dimensional subspace by computing M distinct linear combinations of the variables and utilizing them as predictors for fitting a linear regression model.
  • As model flexibility increases, the test MSE monotonically decreases.
  • How is multicollinearity detected using the variance inflation factor (VIF)?
  • What could be added to a linear regression model to avoid multicollinearity?
  • An estimator is consistent whenever the variance of the estimator approaches zero as the sample size goes to infinity.
  • True or false: For the same total number of knots k, the natural cubic spline uses fewer degrees of freedom than the cubic spline.
  • In ridge regression, what happens to squared bias as s increases?
  • Which sequence leads to the asymptotic variance of θ?
  • True or false: larger values of lambda result in greater effective degrees of freedom for the model.
  • Power parameter in the Tweedie family that corresponds to an inverse-Gaussian distribution.
  • Deviance in generalized linear models and the chi-square distribution: deviance follows a chi-square distribution for all models in the exponential family.
  • Which statement about the kth order statistic from Uniform(0, θ) is correct?
  • If an explanatory variable is uncorrelated with all other explanatory variables, the corresponding variance inflation factor would be 1.
  • What is the formula for the degrees of freedom in a chi-square goodness-of-fit test with g groups and p estimated parameters?
  • With 3 original variables, what is the maximum number of principal components that can be extracted?
  • Greedy Algorithm B for optimization problems is described as which approach?
  • For a natural cubic spline, what is the number of degrees of freedom?
  • Main objective of ridge regression?
  • Ridge regression cannot set any coefficient exactly to zero.
  • In least squares LOOCV, the LOOCV error can be computed from a single fitted model using residuals and leverage.
  • When using the quasi-likelihood approach, is the variance-covariance matrix scaled by an extra dispersion parameter?
  • Regarding kernel density estimation: the larger the bandwidth, the smoother the estimated pdf is.
  • Pr(X>x) with a conditional distribution? (Law of Total Probability)
  • Which statement about ridge regression is true?
  • The deviance is useful for testing the significance of explanatory variables in nested models.
  • In PCR, it is common to use only the first few principal components to predict the response.
  • In ridge regression, what happens to variance as the budget parameter s increases?
  • For modeling a binary outcome such as hospitalization, which distribution and link are most appropriate?
  • According to the Central Limit Theorem, what is the limiting distribution of the sampling distribution of the sample mean?
  • Deviance can be used to test the significance of explanatory variables in nested models.
  • Does the quasi-likelihood approach change the coefficient estimates?
  • Regarding a simple linear relationship, if the irreducible error is zero (e = 0), the 95% confidence interval is equal to the 95% prediction interval.
  • Which model uses more parameters for the same number of knots: a linear spline with k knots or a natural cubic spline with k total knots?
  • Does implementing the quasi-likelihood approach to a GLM change the coefficient estimates?
  • Which statement best describes an ergodic Markov chain?
  • In Poisson regression, which statement about the variance-mean relationship is true?
  • In the Tweedie family, p = 2 corresponds to which distribution?
  • The tuning parameter for ridge regression can be selected using cross-validation.
  • Compared with lasso regression, ridge regression is generally harder to interpret because it retains all predictors in the model.
  • In exponential family distributions, canonical form implies a(y) equals y.
  • What is the formula for pseudo R-squared?
  • Using the ILT to price life insurance policies, the lower bound for the number of deaths during a period is given by which expression?
  • Variance refers to the error arising from the assumptions made in the statistical learning tool.
  • Can a plot of Information be used to visually approximate the MLE of theta?
  • When forming a confidence interval for the difference between two means with known variances, which statistic is used?
  • LOOCV tends to overestimate the test error rate in comparison to validation set approach.
  • True or False: Power is the probability of rejecting the null hypothesis, assuming its false.
  • Which term describes a model that achieves adequate predictive performance using the fewest explanatory variables?
  • GAMs allow for non-linear relationships between each predictor variable and the response.
  • How do you standardize a residual?
  • In regression analysis, does an R-squared value of 0 indicate overfitting?
  • What is a good indication of a probability generating function (PGF)?
  • Cayley’s formula gives the number of labeled trees on n vertices. What is that count?
  • Mallows Cp and AIC are proportional to each other; in fact, they are equal.
  • With E(N)=150, Var(N)=100, z=1.96, what is the ILT lower bound for the number of deaths?
  • For theta=0.5, n=4, k=2, what is the numeric value of E when E = theta * sum_{i=k}^n 1/i?
  • When forming a confidence interval for the ratio of two variances, what kind of statistic is used?
  • Natural splines, regression splines, smoothing splines, local regression, polynomial regression, and step functions are all types of models that can be used as building blocks for GAMs.
  • Which statement about ridge regression is true?
  • Some regularization methods can also perform variable selection by estimating coefficients to be precisely zero.
  • In the Tweedie family, p = 3 corresponds to which distribution?
  • PLS identifies new features in a supervised way by relating them to the target variable.
  • In k-fold cross-validation, the model is fitted a total of k times.
  • What is a commonly used model for times to failure (or survival times)?
  • Which statement is true about backward stepwise selection?
  • A biased estimator can be consistent.
  • True or false: a small deviance indicates a poor fit.
  • In the gambler's ruin scenario with total wealth 75, Ben starts with 40 and Allison with 35. Which expression correctly computes the expected final wealth of Ben?
  • In a regression model that includes an intercept, the sum of residuals is:
  • Which property ensures a Markov chain has a unique stationary distribution and convergence from any starting state?
  • Best subset selection requires fitting all possible subset models, a total of 2^p models.
  • Var(S) for S = sum_{i=1}^N X_i with N ~ Poisson(lambda) and i.i.d. X_i is equal to lambda * E[X^2].
  • Collinearity can exist among three variables even if no single pair shows a high correlation.
  • With least squares regression on a dataset with n observations, the LOOCV estimate for the test MSE can be calculated by fitting a model once.
  • Power parameter in the Tweedie family that corresponds to a compound Poisson-Gamma distribution (1<p<2).
  • Ordinal variables are a type of continuous explanatory variable.
  • To form a confidence interval for the ratio of variances between two populations, which statistic is used?
  • How many minimal cut sets are there for a random graph with n nodes?
  • What is the likelihood ratio critical region for testing H0 against H1?
  • Cp, AIC, BIC, and adjusted R-squared are used to adjust which type of error when evaluating models for model size?
  • When forming a confidence interval for the difference between two means with unknown variances, which statistic is used?
  • As more variables are added to a linear model, the training MSE generally does what?
  • Simon uses a statistical learning method to estimate the number of ears of corn produced per acre. He applies the same method to multiple training data sets and results are similar but not identical. What best describes this method?
  • A high leverage point is an outlier.
  • In kernel density estimation using a Gaussian kernel, the width of the neighborhood is infinite.
  • True or false: Ridge regression is less flexible than OLS and thus results in an improved prediction accuracy when its increase in squared bias is less than its decrease in variance.
  • Using all possible PCs provides the best understanding of the data.
  • True or false: all states in an irreducible Markov chain are recurrent.
  • What is the canonical link for the inverse Gaussian distribution?
  • In the Tweedie family, p = 1 corresponds to which distribution?
  • For a sample from an inverse Gaussian distribution, the MVUE of the mean parameter is:
  • If Y is a complete sufficient statistic for theta and g(Y) is an unbiased estimator of theta, then g(Y) is the MVUE and has the smallest possible variance among all unbiased estimators.
  • Which of the following is NOT listed as a potential cause of unreliable mean squared error estimates?
  • Is a cluster analysis an example of supervised or unsupervised learning?
  • PCA provides low-dimensional linear surfaces that are closest to the observations.
  • To determine if a function should be used as a link function for a GLM, check if the function is monotone and differentiable.
  • If Y is complete sufficient for theta and g(Y) is unbiased for theta, then g(Y) is the MVUE with the smallest variance.
  • In a 5-state Markov chain with two classes {0,1,3} and {2,4}, at least one of the two classes must be recurrent.
  • X ~ Exp(theta). any loss over 10,000 will result in a claim payment of only 10,000 due to policy limits. you observe 4 claim payments: 1000, 3100, 7500, 10000. how would you calculate L(theta)?
  • As lambda increases towards infinity, the ridge penalty term has no effect and the estimates become unconstrained.
  • Which process is characterized as a counting process with integer-valued counts?
  • Which statement about deviance is correct?
  • A logit model gives numerical results that are quite similar to those given by the probit model.
  • Using a canonical link function in a GLM, are the estimates unbiased or biased?
  • The logit model is appropriate when the response variable is binary.
  • If Ti is the time of the ith event, Pr(T2 > 3) represents what probability?
  • In PCR, is it recommended to standardize each predictor prior to generating principal components?
  • Is PCA an example of supervised or unsupervised learning?
  • For a Negative Binomial distribution with parameter r, which expression is the MVUE of the scale parameter?
  • Ridge regression outperforms lasso when the response is a function of many predictors, all with coefficients of roughly equal size.
  • Which method is used to measure the accuracy of a parameter estimate?
  • In a Poisson process, the waiting time until the first event is exponentially distributed with a parameter theta equal to what value in terms of lambda?
  • What's the formula for Cov(X,Y)?
  • When selecting the optimal model, which criterion is preferred?
  • In the formula f_{2,4} = (s_{2,4} - δ_{2,4}) / s_{4,4} used to compute a probability, what does δ represent?
  • A gamma random variable with alpha = 2 and theta = 1 to be when simulating random variables?
  • Greedy Algorithm A is commonly used for which class of optimization problems?
  • A parallel system functions as long as one of the components functions.
  • In simple linear regression, which interval estimates E(Y|X)?
  • True or false: The critical region of a hypothesis test is determined by the significance level and not by the sample observations.
  • What does the likelihood ratio test (LRT) test?
  • K-fold validation has variance reduction compared to LOOCV.
  • When comparing two means assuming known variances for both populations, which statistic is used to form the confidence interval?
  • In a hierarchical normal model where X|θ ~ N(θ, 100^2) and θ ~ N(800, 50^2), what is E[X]?
  • For a binary response variable with a continuous explanatory variable, logistic regression is inappropriate.
  • For a cubic spline with one knot, how can we connect the two pieces of the equation to find coefficient estimates?
  • In least squares regression, the LOOCV estimate for the test MSE can be calculated by fitting a model once.
  • The incremental variance explained by adding another principal component decreases as more components are included.
  • The smallest possible value of a leverage is 0.
  • In Lasso regression, as the regularization parameter lambda increases, what happens to the number of predictors selected?
  • True or false: if all states in a finite Markov chain are recurrent, the Markov chain is irreducible.
  • What is the primary purpose of fitting a saturated model in generalized linear models?
  • PCA selects low-dimensional linear surfaces to maximize the captured variance.
  • N(t) doesn't have to be an integer.
  • Overfitting causes the training error to underestimate the test error.
  • Bias refers to the error arising from the method's sensitivity towards the training data set.
  • In curtate life expectancy, ex_(curtate) relates to survival probability p_x and the life expectancy at age x+1 by which expression?
  • In simple linear regression, does the choice of explanatory variable x affect the total sum of squares?
  • In the context of regression, the statement 'the sample correlation between x and y is equal to the coefficient of determination' is
  • K-fold validation has a computational advantage over LOOCV when k < n.
  • Which statement about the Neyman-Pearson lemma for simple hypotheses is correct?
  • The sum of the leverages across all observations must equal the number of explanatory variables.
  • Is boosting an example of supervised or unsupervised learning?
  • In actuarial notation, A_x denotes the present value of a 1-unit death benefit payable at the end of the year of death for a life aged x.
  • A counting process possesses independent increments if the number of events between s and t is independent of the number between t and t+u for all u>0.
  • Which expression gives the estimated variance of beta_hat in OLS regression?
  • A predictor uncorrelated with others has VIF equal to 1.
  • A logit transformation helps in reducing heteroscedasticity.
  • If the number of PLS components equals the number of predictors in OLS, the forecasted values from both methods are what?
  • In a local regression model, increasing the span s will typically produce what effect on the fitted curve?
  • Minimal cut sets must have at least one component from each minimal path set.
  • What penalty term does lasso regression use?
  • If the branching process starts with n individuals, the extinction probability is π0^n.
  • The sum of leverages across observations equals p+1.
  • PCA finds a low dimension representation of a dataset that contains as much variation as possible.
  • Regression through the origin occurs when the intercept term in the linear equation linking the explanatory variables to the dependent variable is zero or is left out of the equation. Which of the following is true?
  • The cumulative proportion of variance explained increases as more PCs are added.
  • Which of the following expresses ridge regression as a constrained optimization problem?
  • In binomial data, overdispersion manifests as the observed variance exceeding the binomial variance, which is expressed as which formula?
  • Which kernel density estimator distributes mass uniformly in the neighborhood?
  • In quasi-likelihood, how is the variance-covariance matrix adjusted?
  • Power parameter in the Tweedie family that corresponds to a gamma/exponential distribution.
  • In regression context, if there is no linear relationship between x and y, the scatterplot will typically show a random pattern.
  • In actuarial notation, the symbol a_double_dot_40 denotes:
  • How many of the modeling techniques perform dimension reduction: lasso, PLS, PCA, ridge?
  • Which statement best describes how model flexibility affects variance and bias?
  • In a Poisson regression model, what is the offset term?
  • PCA serves as a tool for data visualization.
  • If theta-hat is unbiased and efficient, then theta-hat is the MVUE.
  • True or false: A small value of span s results in a global fit for a local regression.
  • In smoothing spline models, increasing lambda affects the bias-variance tradeoff. Which statement is correct?
  • Which statement about bootstrapping and cross-validation is false?
  • PLS identifies new features in an unsupervised way by approximating the original predictors, similar to PCA.
  • In nested models, the deviance is useful for testing the significance of explanatory variables.
  • Both ridge and lasso regression are regularized methods.
  • What distribution do you use when calculating a confidence interval for beta coefficients?
  • A uniformly MVUE is an estimator such that no other estimator has a smaller variance.
  • In a 10-state Markov chain, which property ensures that all states communicate?
  • For two independent exponential random variables X and Y, what is E[X | X < Y]?
  • For an unbiased estimator, the MSE is always equal to the variance.
  • Ridge regression coefficients are not scale equivariant.
  • True or false: as lambda increases from 0 to infinity, the effective degrees of freedom decrease from n to 2.
  • A consistent estimator cannot be biased.
  • K-fold cross validation requires fitting a model for a total of k times.
  • In a linear regression model, the leverage for each observation is guaranteed to lie between 1/n and 1.
  • In smoothing splines, setting lambda to zero corresponds to no penalty for roughness.
  • what is the excel equation for solving for the cdf in gaussian kernel estimation?
  • A distribution with support depending on theta cannot be a member of the standard exponential family.
  • For the acceptance-rejection method, what's the inequality for f(y)/ (c g(y)) relative to a Uniform(0,1) random variable U?
  • Which statement about leverages in a linear model with intercept and p explanatory variables is correct?
  • If a simple linear model for a probability predicts values outside the [0,1] interval, what is the standard corrective approach?
  • How many minimal path sets are there for a random graph with n nodes, according to the given material?
  • In the given model, E[X] equals E[θ].
  • The deviance is defined as a measure of distance between saturated and fitted model.
  • The first principal component direction of the data is the axis along which the observations vary the most.
  • A biased estimator can only be inconsistent.
  • True or false: Mallows Cp and AIC are proportional to each other.
  • When forming a confidence interval for a population variance, what kind of statistic is used?
  • In a lifetime model where lifetimes are i.i.d. exponential with mean θ, the expected value of the k-th order statistic X_(k) is equal to which expression?
  • In ordinary least squares, the sum of residuals equals which value?
  • When applying the quasi-likelihood approach to a GLM, does it change the point estimates of the coefficients?
  • True or false: Local regression should not be used in a high-dimensional setting.
  • If theta-hat is the MVUE, theta-hat is efficient.
  • n^(n-2) is associated with the count of which standard combinatorial object?
  • Kernel density estimation is used to estimate which component of a distribution?
  • In PCA, the first few principal components are often sufficient to get a good understanding of the data.
  • A logit model applies when the response variable counts the number of events occurring.
  • The efficiency of an estimator is defined as the Rao-Cramer lower bound divided by the estimator's variance.
  • Which distribution is commonly used to model life data due to its flexible hazard function?
  • If all regression errors are identically zero, what is the R-squared value?
  • In smoothing spline models, increasing lambda has what effect on the bias-variance tradeoff?
  • In life contingencies notation, Ax denotes the present value of what kind of benefit?
  • Which statement is true about the scale behavior of ridge regression?
  • Is an absorbing state considered transient or recurrent?
  • Ridge regression objective can be formulated as minimizing SSR subject to which constraint?
  • Positive recurrence is a class property; if state 2 is positive recurrent, then state 4 must be positive recurrent.
  • N(t) must be greater than or equal to 0 is a property of counting processes.
  • True or false: The model containing all predictors will always have the smallest residual sum of squares and largest R-squared.
  • Ridge regression tends to shrink coefficient estimates toward zero and typically does not set any coefficients exactly to zero.
  • To determine asymptotic unbiasedness, which condition must hold?
  • Mallows Cp involves SSE, p, and MSEfull.
  • Which statistic is used to construct a confidence interval for a single population variance?
  • Relative to the least squares estimates, shrinkage has the effect of reducing bias.
  • True or False: The F-statistic used to test a predictor in regression is computed as the ratio of its mean square to the mean square error from the model with the most predictors.
  • Under the saturated model, what is the predicted value for each observation?
  • In ridge regression, which statement is true?
  • Forward stepwise selection requires fitting 1 + (1/2)(p*(p+1)) models.
  • Power parameter in the Tweedie family that corresponds to a normal distribution.
  • For a Bernoulli response in a generalized linear model, which set of link functions can be used?
  • In the Tweedie family, p in (1,2) corresponds to which distribution?
  • Do the beta_hat values of a ridge regression procedure provide unbiased estimators of the corresponding beta model parameters?
  • Given that the corresponding coefficient estimates for all models with two explanatory variables are the same, which link function will produce a prediction for an observation that is always the greatest?
  • In a parallel system, the system fails only when all components fail.
  • Which inequality defines the likelihood ratio test critical region?
  • Which kernel density estimation fact is true?
  • What is the test statistic for testing the significance of a single parameter in a regression model?
  • It is possible to directly estimate a model's test error using a validation set or cross-validation.
  • True or false: We should choose a model with a low training error when selecting the optimal model.
  • What is the matrix S defined as in the method for transient absorption probabilities?
  • Backward stepwise selection cannot be used when the dataset has as many predictors as observations (n < p+1).
  • What is the MVUE of a normal distribution with a defined mean?
  • Using quasi-likelihood, what must be assumed about the relationship between the mean and the variance?
  • The statement that the proportion of variance explained by an additional principal component increases as more PCs are added is true or false?
  • LOOCV requires fitting a model a total of n times.
  • What is the typical statement about removing outliers on model fit?
  • A 5-state Markov chain with two classes {0,1,3} and {2,4} is ergodic.
  • A large value of a leverage indicates the presence of an outlier.
  • For a k-out-of-n system with iid exponential components, the expected lifetime is theta * sum_{i=k}^n 1/i. If theta=2, n=5, k=3, what is the expected lifetime?
  • What happens to training mean squared error as model flexibility increases?
  • Shrinkage reduces variance at the cost of a small increase in bias.
  • Conditional on θ, X follows a Normal distribution with mean θ and variance 100^2.
  • Increasing the significance level would decrease the power of a test.
  • In the context of the material, when estimating an exponential distribution, the MLE of the mean equals the sample mean.
  • Dimension reduction reduces the number of predictors by projecting onto a lower-dimensional space.
  • True or false about a smoothing spline model fit to data using the tuning parameter lambda: larger values of lambda result in smoother splines.
  • A minimal path set is a minimal set of components whose functioning guarantees the functioning of the system.
  • A limitation of GAMs is that interactions cannot be added to the model.
  • What best describes the type of problem Angela is solving by clustering shoppers to target ads?
  • A Markov chain that is irreducible, positive recurrent, and aperiodic is called
  • The maximum possible value of the standard error of the population variance is achieved under which condition?
  • In kernel density estimation, what does the pdf indicate we are looking for?
  • Poisson regression models are capable of handling varying exposure by allowing the exposure term to differ across observations.
  • In a likelihood ratio test, which of the following is a correct statement about the null hypothesis?
  • A Markov chain with two communicating classes is not irreducible.
  • In the gambler's ruin scenario, what is Ben's expected final wealth when the total is 75 and the win probability is 0.5?
  • All collinearity problems can be detected by inspection of the correlation matrix.
  • Which statement is NOT an assumption when using pooled variances for two-sample tests?
  • The MVUE of the scale parameter beta in Gamma(shape=alpha, scale=beta) is which expression?
  • Deviance is a measure used to assess the quality of fit for nested models.
  • Regularized regression methods include ridge and lasso, and their purpose is to prevent overfitting.
  • In the same model, what is Var(X)?
  • When given a CDF of a distribution, how do you obtain the CDF of the kth order statistic Y_(k)?
  • Which expression is the MLE for the exponential distribution when data are censored and truncated?
  • Which of the following describes the characteristics of a non-homogeneous Poisson random variable?
  • The estimated variance of a coefficient in OLS is given by which expression?
  • In the Poisson context, which statement describes overdispersion?
  • PLS is a dimension reduction method.
  • PCR can reduce overfitting.
  • The efficiency of theta-hat is the estimator's variance divided by the Rao-Cramer lower bound.
  • Which statistic would you use to form a confidence interval for the difference of two means when the variances are unknown?
  • The statement 'The logit link corresponds to a logistic distribution, and the probit link corresponds to a standard normal distribution' is true.
  • What's the formula for Var(X+Y)?
  • In simple linear regression, the F-statistic for the model equals the square of the t-statistic for the slope parameter.
  • When forming a confidence interval for a mean with a known variance, which statistic is used in the calculation?
  • In testing whether a source is significant, the test statistic is the mean square of that source divided by the MSE of the model that has the most predictors.
  • If X ~ Exp(λx) and Y ~ Exp(λy) are independent, what is E[min(X,Y)]?
  • A uniformly MVUE is defined as an estimator such that no other estimator has a smaller variance.
  • R^2 is the ratio of the regression sum of squares to the total sum of squares.
  • Before applying ridge regression, predictors should be standardized because ridge is not scale invariant.
  • Which of the following statements about the prediction interval and the range it measures is true?
  • If the stationary probability of state 0 in a Markov chain is 0.2, what is the expected return time to state 0?
  • Overdispersion occurs when the observed variance is larger than the mean for a Poisson model, or when the observed variance is larger than the calculated variance for a binomial model. Which statement best captures overdispersion in common models?
  • What penalty term does ridge regression use?
  • Best subset selection results in a nested set of best models with different numbers of predictors.
  • Ridge regression uses an L2 penalty on coefficients.
  • It is possible to estimate test error by adjusting training error to account for bias due to overfitting.
  • In Lasso regression, as lambda increases, the squared bias of the parameters in the model tends to
  • For Poisson processes, counts in disjoint intervals are independent.
  • Residual sum of squares is monotonic with respect to the number of predictors.
  • Increasing model flexibility decreases variance.
  • In a Markov chain, a state that is guaranteed to be revisited eventually is called a
Subscribe

Get the latest from Examzify

You can unsubscribe at any time. Read our privacy policy