Skip to content
Work Free practice Coding course Blog Method Results Why me About Enquire Book a call

Practice · Quant prep · Statistics & machine learning

Statistics & machine learning interview questions

43 questions with worked solutions. Regression, regularisation, classification metrics, estimation and backtest pitfalls.

26 easy · 17 medium · 0 hard · Question type reported at: Two SigmaCitadelD. E. ShawMillennium

Practise timed: 30 questions · 60s each All Quant prep topics


1 Easy Type asked atTwo SigmaCitadelD. E. Shaw
The unbiased sample variance of nn observations divides the sum of squared deviations by:
  1. nn
  2. n−1n-1
  3. n+1n+1
  4. n\displaystyle \sqrt n
Show answer
Answer: n−1n-1
One degree of freedom is used up estimating the mean (Bessel's correction).
2 Easy Type asked atTwo SigmaCitadelD. E. Shaw
The ordinary least squares estimate in y=Xβ+εy=X\beta+\varepsilon is:
  1. X⊤y\displaystyle X^\top y
  2. (X⊤X)−1X⊤y\displaystyle (X^\top X)^{-1}X^\top y
  3. (XX⊤)−1y\displaystyle (XX^\top)^{-1}y
  4. X−1y\displaystyle X^{-1}y
Show answer
Answer: (X⊤X)−1X⊤y(X^\top X)^{-1}X^\top y
Setting the gradient of ∥y−Xβ∥2\lVert y-X\beta\rVert^2 to zero gives the normal equations X⊤Xβ=X⊤yX^\top X\beta=X^\top y.
3 Easy Type asked atTwo SigmaCitadelD. E. Shaw
In a simple linear regression the correlation between xx and yy is 0.60.6. What is R2R^2?
Show answer
Answer: 0.360.36
With one regressor, R2=ρ2R^2=\rho^2.
4 Medium Type asked atTwo SigmaCitadelD. E. Shaw
In kk-nearest-neighbours, increasing kk generally:
  1. raises variance and lowers bias
  2. raises bias and lowers variance
  3. raises both
  4. has no effect
Show answer
Answer: Raises bias, lowers variance
Averaging over more neighbours smooths the prediction: less sensitive to noise, less able to follow local structure.
5 Easy Type asked atTwo SigmaCitadelD. E. Shaw
Which penalty tends to set some coefficients exactly to zero?
  1. L2 (ridge)
  2. L1 (lasso)
  3. Neither
  4. Both equally
Show answer
Answer: L1 (lasso)
The L1 ball has corners on the axes, so the optimum often lands where some coordinates are 0.
6 Easy Type asked atTwo SigmaCitadelD. E. Shaw
Logistic regression is fitted by minimising:
  1. squared error
  2. cross-entropy (log loss)
  3. hinge loss
  4. absolute error
Show answer
Answer: Cross-entropy
It is maximum likelihood for a Bernoulli model, which is the log loss.
7 Easy Type asked atTwo SigmaCitadelD. E. Shaw
Low training error with much higher test error is a sign of:
  1. underfitting
  2. overfitting
  3. data leakage only
  4. a perfect model
Show answer
Answer: Overfitting
The model has learned noise specific to the training set.
8 Easy Type asked atTwo SigmaCitadelD. E. Shaw
A classifier has 40 true positives and 10 false positives. What is its precision?
Show answer
Answer: 0.80.8
TPTP+FP=4050\dfrac{TP}{TP+FP}=\dfrac{40}{50}.
9 Easy Type asked atTwo SigmaCitadelD. E. Shaw
A classifier has 40 true positives and 60 false negatives. What is its recall?
Show answer
Answer: 0.40.4
TPTP+FN=40100\dfrac{TP}{TP+FN}=\dfrac{40}{100}.
10 Medium Type asked atTwo SigmaCitadelD. E. Shaw
Precision is 0.80.8 and recall is 0.40.4. What is the F1 score?
Show answer
Answer: 815≈0.533\tfrac{8}{15}\approx0.533
The harmonic mean: 2⋅0.8⋅0.40.8+0.4=0.641.2\dfrac{2\cdot0.8\cdot0.4}{0.8+0.4}=\dfrac{0.64}{1.2}.
11 Easy Type asked atTwo SigmaCitadelD. E. Shaw
99% of examples are negative. What accuracy does a model get by always predicting "negative"?
Show answer
Answer: 99%99\%
Which is why accuracy is a poor metric for imbalanced classes; look at precision, recall or AUC.
12 Easy Type asked atTwo SigmaCitadelD. E. Shaw
The first principal component is the direction that:
  1. minimises variance of the projection
  2. maximises variance of the projection
  3. is orthogonal to the data mean
  4. maximises correlation with the target
Show answer
Answer: Maximises projected variance
It is the top eigenvector of the covariance matrix.
13 Medium Type asked atTwo SigmaCitadelD. E. Shaw
A covariance matrix has eigenvalues 5,3,25,3,2. What fraction of variance does the first principal component explain?
Show answer
Answer: 0.50.5
55+3+2\dfrac{5}{5+3+2}.
14 Easy Type asked atTwo SigmaCitadelD. E. Shaw
A population has standard deviation 10. What is the standard error of the mean of 100 observations?
Show answer
Answer: 11
σ/n=10/10\sigma/\sqrt n=10/10.
15 Medium Type asked atTwo SigmaCitadelD. E. Shaw
With σ=20\sigma=20 and n=400n=400, what is the half-width of a 95% confidence interval for the mean?
Show answer
Answer: 1.961.96
1.96⋅σ/n=1.96⋅20/201.96\cdot\sigma/\sqrt n=1.96\cdot20/20.
16 Easy Type asked atTwo SigmaCitadelD. E. Shaw
If the learning rate in gradient descent is too large, the loss typically:
  1. converges faster
  2. oscillates or diverges
  3. reaches the global minimum
  4. stays constant
Show answer
Answer: Oscillates or diverges
Steps overshoot the minimum; beyond a threshold (about 2/L2/L for an LL-smooth loss) the iterates blow up.
17 Medium Type asked atTwo SigmaCitadelD. E. Shaw
What is the derivative of the sigmoid σ(x)=11+e−x\sigma(x)=\dfrac1{1+e^{-x}} at x=0x=0?
Show answer
Answer: 14\tfrac14
σ′(x)=σ(x)(1−σ(x))=12⋅12\sigma'(x)=\sigma(x)(1-\sigma(x))=\tfrac12\cdot\tfrac12.
18 Easy Type asked atTwo SigmaCitadelD. E. Shaw
Which activation most directly mitigates vanishing gradients in deep networks?
  1. sigmoid
  2. tanh
  3. ReLU
  4. step function
Show answer
Answer: ReLU
Its gradient is 1 for positive inputs, so it does not shrink through many layers the way sigmoid and tanh do.
19 Easy Type asked atTwo SigmaCitadelD. E. Shaw
The main purpose of kk-fold cross-validation is to:
  1. speed up training
  2. estimate out-of-sample performance
  3. increase the training set
  4. remove outliers
Show answer
Answer: Estimate out-of-sample performance
Each fold is held out once, giving a less noisy estimate than a single split.
20 Medium Type asked atTwo SigmaCitadelD. E. Shaw
A backtest uses each day's closing price to decide a trade executed at that same close. This is an example of:
  1. survivorship bias
  2. look-ahead bias
  3. regularisation
  4. mean reversion
Show answer
Answer: Look-ahead bias
The signal uses information not available when the trade would actually have been placed.
21 Medium Type asked atTwo SigmaCitadelD. E. Shaw
Strong multicollinearity among regressors mainly causes:
  1. biased coefficients
  2. inflated variance of the coefficients
  3. a lower R2\displaystyle R^2
  4. non-normal residuals
Show answer
Answer: Inflated coefficient variance
OLS stays unbiased, but (X⊤X)−1(X^\top X)^{-1} becomes large, so estimates are unstable.
22 Easy Type asked atTwo SigmaCitadelD. E. Shaw
Dropout in neural networks is primarily a form of:
  1. optimisation
  2. regularisation
  3. normalisation
  4. data augmentation
Show answer
Answer: Regularisation
Randomly removing units stops the network relying on any single pathway.
23 Easy Type asked atTwo SigmaCitadelD. E. Shaw
Naive Bayes assumes the features are:
  1. normally distributed
  2. conditionally independent given the class
  3. uncorrelated with the class
  4. identically distributed
Show answer
Answer: Conditionally independent given the class
That is what makes the likelihood factorise into a product.
24 Easy Type asked atTwo SigmaCitadelD. E. Shaw
What is the entropy of a fair coin, in bits?
Show answer
Answer: 11
−2⋅12log⁡212=1-2\cdot\tfrac12\log_2\tfrac12=1.
25 Medium Type asked atTwo SigmaCitadelD. E. Shaw
What is the entropy of a fair six-sided die, in bits?
Show answer
Answer: log⁡26≈2.585\log_26\approx2.585
A uniform distribution on nn outcomes has entropy log⁡2n\log_2n.
26 Easy Type asked atTwo SigmaCitadelD. E. Shaw
Kullback–Leibler divergence DKL(P∥Q)D_{KL}(P\Vert Q) is:
  1. symmetric in PP and QQ
  2. not symmetric, and always ≥0\ge0
  3. sometimes negative
  4. a true metric
Show answer
Answer: Not symmetric, always ≥ 0
Gibbs' inequality gives non-negativity; in general DKL(P∥Q)≠DKL(Q∥P)D_{KL}(P\Vert Q)\ne D_{KL}(Q\Vert P).
27 Medium Type asked atTwo SigmaCitadelD. E. Shaw
X∼N(0,1)X\sim N(0,1). What is corr⁡(X,X2)\operatorname{corr}(X,X^2)?
Show answer
Answer: 00
cov⁡(X,X2)=E[X3]=0\operatorname{cov}(X,X^2)=\mathbb E[X^3]=0. Zero correlation despite complete dependence.
28 Easy Type asked atTwo SigmaCitadelD. E. Shaw
A coin gives 7 heads in 10 tosses. What is the maximum-likelihood estimate of P(heads)P(\text{heads})?
Show answer
Answer: 0.70.7
For a Bernoulli model the MLE is the sample proportion.
29 Medium Type asked atTwo SigmaCitadelD. E. Shaw
Waiting times are modelled as exponential and the sample mean is 4. What is the MLE of the rate λ\lambda?
Show answer
Answer: 0.250.25
λ^=1/xˉ\hat\lambda=1/\bar x.
30 Medium Type asked atTwo SigmaCitadelD. E. Shaw
A strategy has a daily Sharpe ratio of 0.10.1. Roughly what is its annualised Sharpe ratio (252 trading days)?
Show answer
Answer: ≈1.59\approx1.59
Mean scales with TT, volatility with T\sqrt T, so the Sharpe ratio scales with 252≈15.9\sqrt{252}\approx15.9.
31 Medium Type asked atTwo SigmaCitadelD. E. Shaw
Two uncorrelated assets each have 20% volatility. What is the volatility of an equal-weight portfolio, in %?
Show answer
Answer: ≈14.1%\approx14.1\%
0.52⋅0.22⋅2=0.2/2\sqrt{0.5^2\cdot0.2^2\cdot2}=0.2/\sqrt2.
32 Easy Type asked atTwo SigmaCitadelD. E. Shaw
You test 20 strategies that truly have no edge, each at the 5% significance level. How many false positives do you expect?
Show answer
Answer: 11
20×0.0520\times0.05. The reason backtest results need multiple-testing corrections.
33 Easy Type asked atTwo SigmaCitadelD. E. Shaw
With a Bonferroni correction for 20 tests at overall level 0.05, what threshold should each test use?
Show answer
Answer: 0.00250.0025
0.05/200.05/20.
34 Easy Type asked atTwo SigmaCitadelD. E. Shaw
That the distribution of a sample mean is approximately normal for large nn is:
  1. the law of large numbers
  2. the central limit theorem
  3. Chebyshev's inequality
  4. Jensen's inequality
Show answer
Answer: The central limit theorem
The law of large numbers says the mean converges; the CLT describes its fluctuations.
35 Medium Type asked atTwo SigmaCitadelD. E. Shaw
The ridge regression estimate with penalty λ∥β∥2\lambda\lVert\beta\rVert^2 is:
  1. (X⊤X)−1X⊤y\displaystyle (X^\top X)^{-1}X^\top y
  2. (X⊤X+λI)−1X⊤y\displaystyle (X^\top X+\lambda I)^{-1}X^\top y
  3. (X⊤X−λI)−1X⊤y\displaystyle (X^\top X-\lambda I)^{-1}X^\top y
  4. λ(X⊤X)−1X⊤y\displaystyle \lambda(X^\top X)^{-1}X^\top y
Show answer
Answer: (X⊤X+λI)−1X⊤y(X^\top X+\lambda I)^{-1}X^\top y
The penalty adds 2λβ2\lambda\beta to the gradient, shifting the normal equations by λI\lambda I.
36 Easy Type asked atTwo SigmaCitadelD. E. Shaw
What is the training error of 1-nearest-neighbour classification (no duplicate points)?
Show answer
Answer: 00
Each training point is its own nearest neighbour. A reminder that training error says little.
37 Medium Type asked atTwo SigmaCitadelD. E. Shaw
You fit ordinary least squares with 100 features and only 50 observations. What happens?
  1. The fit is unique and unbiased
  2. There are infinitely many coefficient vectors that fit the data exactly
  3. The residual sum of squares is minimised at a unique interior point
  4. The estimator is consistent
  5. The variance of the estimator is zero
Show answer
Answer: B. There are infinitely many coefficient vectors that fit the data exactly
With more features than observations the design matrix has a non-trivial null space, so the normal equations are underdetermined: infinitely many solutions give a perfect in-sample fit. This is why regularisation is needed.
38 Easy Type asked atTwo SigmaCitadelD. E. Shaw
What is the area under the ROC curve for a classifier that assigns scores at random?
Show answer
Answer: 0.50.5
Random scores rank a positive above a negative half the time, and AUC is exactly that probability.
39 Easy Type asked atTwo SigmaCitadelD. E. Shaw
In ridge regression, what happens to the coefficients as the penalty λ→∞\lambda\to\infty?
  1. They approach the OLS estimates
  2. They approach zero
  3. They diverge
  4. They approach one
  5. They become exactly sparse
Show answer
Answer: B. They approach zero
The penalty dominates the loss, so the minimiser shrinks to the zero vector. Ridge shrinks smoothly; lasso is the one that sets coefficients exactly to zero.
40 Medium Type asked atTwo SigmaCitadelD. E. Shaw
In kk-nearest-neighbour regression, what happens as kk increases?
  1. Bias falls and variance rises
  2. Bias rises and variance falls
  3. Both rise
  4. Both fall
  5. Neither changes
Show answer
Answer: B. Bias rises and variance falls
Averaging over more neighbours smooths the fit: less sensitive to noise (lower variance) but less able to follow the true function (higher bias).
41 Medium Type asked atTwo SigmaCitadelD. E. Shaw
A classifier has precision 0.6 and recall 0.3. What is its F1F_1 score?
Show answer
Answer: 0.40.4
F1F_1 is the harmonic mean: 2(0.6)(0.3)0.6+0.3=0.360.9\tfrac{2(0.6)(0.3)}{0.6+0.3}=\tfrac{0.36}{0.9}.
42 Easy Type asked atTwo SigmaCitadelD. E. Shaw
What is the correlation between XX and 2X+32X+3?
Show answer
Answer: 11
Correlation is invariant to positive linear transformations.
43 Medium Type asked atTwo SigmaCitadelD. E. Shaw
In a simple regression of yy on xx, the correlation is 0.6 and the two standard deviations are equal. What is the slope?
Show answer
Answer: 0.60.6
β^=ρ sysx\hat\beta=\rho\,\tfrac{s_y}{s_x}, and the ratio is 1. This is regression to the mean: the slope is below 1 even though the variables are equally spread.

Keep practising

Want someone to work through these with you? Quant interview preparation, one to one, or book a free 20-minute call.