The unbiased sample variance of n observations divides the sum of squared deviations by:
n
n−1
n+1
n
Show answer
Answer:n−1
One degree of freedom is used up estimating the mean (Bessel's correction).
2EasyType asked atTwo SigmaCitadelD. E. Shaw
The ordinary least squares estimate in y=Xβ+ε is:
X⊤y
(X⊤X)−1X⊤y
(XX⊤)−1y
X−1y
Show answer
Answer:(X⊤X)−1X⊤y
Setting the gradient of ∥y−Xβ∥2 to zero gives the normal equations X⊤Xβ=X⊤y.
3EasyType asked atTwo SigmaCitadelD. E. Shaw
In a simple linear regression the correlation between x and y is 0.6. What is R2?
Show answer
Answer:0.36
With one regressor, R2=ρ2.
4MediumType asked atTwo SigmaCitadelD. E. Shaw
In k-nearest-neighbours, increasing k generally:
raises variance and lowers bias
raises bias and lowers variance
raises both
has no effect
Show answer
Answer: Raises bias, lowers variance
Averaging over more neighbours smooths the prediction: less sensitive to noise, less able to follow local structure.
5EasyType asked atTwo SigmaCitadelD. E. Shaw
Which penalty tends to set some coefficients exactly to zero?
L2 (ridge)
L1 (lasso)
Neither
Both equally
Show answer
Answer: L1 (lasso)
The L1 ball has corners on the axes, so the optimum often lands where some coordinates are 0.
6EasyType asked atTwo SigmaCitadelD. E. Shaw
Logistic regression is fitted by minimising:
squared error
cross-entropy (log loss)
hinge loss
absolute error
Show answer
Answer: Cross-entropy
It is maximum likelihood for a Bernoulli model, which is the log loss.
7EasyType asked atTwo SigmaCitadelD. E. Shaw
Low training error with much higher test error is a sign of:
underfitting
overfitting
data leakage only
a perfect model
Show answer
Answer: Overfitting
The model has learned noise specific to the training set.
8EasyType asked atTwo SigmaCitadelD. E. Shaw
A classifier has 40 true positives and 10 false positives. What is its precision?
Show answer
Answer:0.8
TP+FPTP=5040.
9EasyType asked atTwo SigmaCitadelD. E. Shaw
A classifier has 40 true positives and 60 false negatives. What is its recall?
Show answer
Answer:0.4
TP+FNTP=10040.
10MediumType asked atTwo SigmaCitadelD. E. Shaw
Precision is 0.8 and recall is 0.4. What is the F1 score?
Show answer
Answer:158≈0.533
The harmonic mean: 0.8+0.42⋅0.8⋅0.4=1.20.64.
11EasyType asked atTwo SigmaCitadelD. E. Shaw
99% of examples are negative. What accuracy does a model get by always predicting "negative"?
Show answer
Answer:99%
Which is why accuracy is a poor metric for imbalanced classes; look at precision, recall or AUC.
12EasyType asked atTwo SigmaCitadelD. E. Shaw
The first principal component is the direction that:
minimises variance of the projection
maximises variance of the projection
is orthogonal to the data mean
maximises correlation with the target
Show answer
Answer: Maximises projected variance
It is the top eigenvector of the covariance matrix.
13MediumType asked atTwo SigmaCitadelD. E. Shaw
A covariance matrix has eigenvalues 5,3,2. What fraction of variance does the first principal component explain?
Show answer
Answer:0.5
5+3+25.
14EasyType asked atTwo SigmaCitadelD. E. Shaw
A population has standard deviation 10. What is the standard error of the mean of 100 observations?
Show answer
Answer:1
σ/n=10/10.
15MediumType asked atTwo SigmaCitadelD. E. Shaw
With σ=20 and n=400, what is the half-width of a 95% confidence interval for the mean?
Show answer
Answer:1.96
1.96⋅σ/n=1.96⋅20/20.
16EasyType asked atTwo SigmaCitadelD. E. Shaw
If the learning rate in gradient descent is too large, the loss typically:
converges faster
oscillates or diverges
reaches the global minimum
stays constant
Show answer
Answer: Oscillates or diverges
Steps overshoot the minimum; beyond a threshold (about 2/L for an L-smooth loss) the iterates blow up.
17MediumType asked atTwo SigmaCitadelD. E. Shaw
What is the derivative of the sigmoid σ(x)=1+e−x1 at x=0?
Show answer
Answer:41
σ′(x)=σ(x)(1−σ(x))=21⋅21.
18EasyType asked atTwo SigmaCitadelD. E. Shaw
Which activation most directly mitigates vanishing gradients in deep networks?
sigmoid
tanh
ReLU
step function
Show answer
Answer: ReLU
Its gradient is 1 for positive inputs, so it does not shrink through many layers the way sigmoid and tanh do.
19EasyType asked atTwo SigmaCitadelD. E. Shaw
The main purpose of k-fold cross-validation is to:
speed up training
estimate out-of-sample performance
increase the training set
remove outliers
Show answer
Answer: Estimate out-of-sample performance
Each fold is held out once, giving a less noisy estimate than a single split.
20MediumType asked atTwo SigmaCitadelD. E. Shaw
A backtest uses each day's closing price to decide a trade executed at that same close. This is an example of:
survivorship bias
look-ahead bias
regularisation
mean reversion
Show answer
Answer: Look-ahead bias
The signal uses information not available when the trade would actually have been placed.
21MediumType asked atTwo SigmaCitadelD. E. Shaw
Strong multicollinearity among regressors mainly causes:
biased coefficients
inflated variance of the coefficients
a lower R2
non-normal residuals
Show answer
Answer: Inflated coefficient variance
OLS stays unbiased, but (X⊤X)−1 becomes large, so estimates are unstable.
22EasyType asked atTwo SigmaCitadelD. E. Shaw
Dropout in neural networks is primarily a form of:
optimisation
regularisation
normalisation
data augmentation
Show answer
Answer: Regularisation
Randomly removing units stops the network relying on any single pathway.
23EasyType asked atTwo SigmaCitadelD. E. Shaw
Naive Bayes assumes the features are:
normally distributed
conditionally independent given the class
uncorrelated with the class
identically distributed
Show answer
Answer: Conditionally independent given the class
That is what makes the likelihood factorise into a product.
24EasyType asked atTwo SigmaCitadelD. E. Shaw
What is the entropy of a fair coin, in bits?
Show answer
Answer:1
−2⋅21log221=1.
25MediumType asked atTwo SigmaCitadelD. E. Shaw
What is the entropy of a fair six-sided die, in bits?
Show answer
Answer:log26≈2.585
A uniform distribution on n outcomes has entropy log2n.
26EasyType asked atTwo SigmaCitadelD. E. Shaw
Kullback–Leibler divergence DKL(P∥Q) is:
symmetric in P and Q
not symmetric, and always ≥0
sometimes negative
a true metric
Show answer
Answer: Not symmetric, always ≥ 0
Gibbs' inequality gives non-negativity; in general DKL(P∥Q)=DKL(Q∥P).
27MediumType asked atTwo SigmaCitadelD. E. Shaw
X∼N(0,1). What is corr(X,X2)?
Show answer
Answer:0
cov(X,X2)=E[X3]=0. Zero correlation despite complete dependence.
28EasyType asked atTwo SigmaCitadelD. E. Shaw
A coin gives 7 heads in 10 tosses. What is the maximum-likelihood estimate of P(heads)?
Show answer
Answer:0.7
For a Bernoulli model the MLE is the sample proportion.
29MediumType asked atTwo SigmaCitadelD. E. Shaw
Waiting times are modelled as exponential and the sample mean is 4. What is the MLE of the rate λ?
Show answer
Answer:0.25
λ^=1/xˉ.
30MediumType asked atTwo SigmaCitadelD. E. Shaw
A strategy has a daily Sharpe ratio of 0.1. Roughly what is its annualised Sharpe ratio (252 trading days)?
Show answer
Answer:≈1.59
Mean scales with T, volatility with T, so the Sharpe ratio scales with 252≈15.9.
31MediumType asked atTwo SigmaCitadelD. E. Shaw
Two uncorrelated assets each have 20% volatility. What is the volatility of an equal-weight portfolio, in %?
Show answer
Answer:≈14.1%
0.52⋅0.22⋅2=0.2/2.
32EasyType asked atTwo SigmaCitadelD. E. Shaw
You test 20 strategies that truly have no edge, each at the 5% significance level. How many false positives do you expect?
Show answer
Answer:1
20×0.05. The reason backtest results need multiple-testing corrections.
33EasyType asked atTwo SigmaCitadelD. E. Shaw
With a Bonferroni correction for 20 tests at overall level 0.05, what threshold should each test use?
Show answer
Answer:0.0025
0.05/20.
34EasyType asked atTwo SigmaCitadelD. E. Shaw
That the distribution of a sample mean is approximately normal for large n is:
the law of large numbers
the central limit theorem
Chebyshev's inequality
Jensen's inequality
Show answer
Answer: The central limit theorem
The law of large numbers says the mean converges; the CLT describes its fluctuations.
35MediumType asked atTwo SigmaCitadelD. E. Shaw
The ridge regression estimate with penalty λ∥β∥2 is:
(X⊤X)−1X⊤y
(X⊤X+λI)−1X⊤y
(X⊤X−λI)−1X⊤y
λ(X⊤X)−1X⊤y
Show answer
Answer:(X⊤X+λI)−1X⊤y
The penalty adds 2λβ to the gradient, shifting the normal equations by λI.
36EasyType asked atTwo SigmaCitadelD. E. Shaw
What is the training error of 1-nearest-neighbour classification (no duplicate points)?
Show answer
Answer:0
Each training point is its own nearest neighbour. A reminder that training error says little.
37MediumType asked atTwo SigmaCitadelD. E. Shaw
You fit ordinary least squares with 100 features and only 50 observations. What happens?
The fit is unique and unbiased
There are infinitely many coefficient vectors that fit the data exactly
The residual sum of squares is minimised at a unique interior point
The estimator is consistent
The variance of the estimator is zero
Show answer
Answer: B. There are infinitely many coefficient vectors that fit the data exactly
With more features than observations the design matrix has a non-trivial null space, so the normal equations are underdetermined: infinitely many solutions give a perfect in-sample fit. This is why regularisation is needed.
38EasyType asked atTwo SigmaCitadelD. E. Shaw
What is the area under the ROC curve for a classifier that assigns scores at random?
Show answer
Answer:0.5
Random scores rank a positive above a negative half the time, and AUC is exactly that probability.
39EasyType asked atTwo SigmaCitadelD. E. Shaw
In ridge regression, what happens to the coefficients as the penalty λ→∞?
They approach the OLS estimates
They approach zero
They diverge
They approach one
They become exactly sparse
Show answer
Answer: B. They approach zero
The penalty dominates the loss, so the minimiser shrinks to the zero vector. Ridge shrinks smoothly; lasso is the one that sets coefficients exactly to zero.
40MediumType asked atTwo SigmaCitadelD. E. Shaw
In k-nearest-neighbour regression, what happens as k increases?
Bias falls and variance rises
Bias rises and variance falls
Both rise
Both fall
Neither changes
Show answer
Answer: B. Bias rises and variance falls
Averaging over more neighbours smooths the fit: less sensitive to noise (lower variance) but less able to follow the true function (higher bias).
41MediumType asked atTwo SigmaCitadelD. E. Shaw
A classifier has precision 0.6 and recall 0.3. What is its F1 score?
Show answer
Answer:0.4
F1 is the harmonic mean: 0.6+0.32(0.6)(0.3)=0.90.36.
42EasyType asked atTwo SigmaCitadelD. E. Shaw
What is the correlation between X and 2X+3?
Show answer
Answer:1
Correlation is invariant to positive linear transformations.
43MediumType asked atTwo SigmaCitadelD. E. Shaw
In a simple regression of y on x, the correlation is 0.6 and the two standard deviations are equal. What is the slope?
Show answer
Answer:0.6
β^=ρsxsy, and the ratio is 1. This is regression to the mean: the slope is below 1 even though the variables are equally spread.