Data Science & Applied Analytics Pod
Deep dive into statistical inference, SQL window functions, exploratory data analysis with Pandas, and scikit-learn models.
Take turns: One student acts as interviewer while the other answers using the STAR technique (Situation, Task, Action, Result).
What is the difference between RANK(), DENSE_RANK(), and ROW_NUMBER() in SQL window functions?
Detail handling of duplicate tie values: ROW_NUMBER increments sequentially, RANK leaves gaps, DENSE_RANK does not leave gaps.
How do you detect and handle data leakage when preparing train and test splits for supervised learning?
Explain why scaling and imputation parameters must fit ONLY on the training split before transforming the test split.
Can you explain Simpson's Paradox using an intuitive real-world data scenario?
Provide an example where a trend appears in different groups of data but disappears or reverses when the groups are combined.
When would you choose an ROC-AUC score over raw accuracy for evaluating binary classification?
Discuss class imbalance (e.g. 99% negative vs 1% fraud detection) where 99% accuracy is completely deceptive.
Describe how you communicate complex statistical findings to non-technical business stakeholders.
Explain translating technical metrics into business impact, risk reduction, and actionable next steps.