What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
The source behind the title “25 Questions to Detect Fake Data Scientists” contains 20 interview prompts, not 25. Andrew Fogg’s list, published by KDnuggets on January 1, 2016, covers modeling, statistics, experimental design, data handling, and communication. It does not establish a 25-question version or prove that any answer predicts job performance. Use the prompts below to explore a candidate’s reasoning and work—not to label someone “fake.”
How to use these questions fairly
Data science draws on several disciplines. A candidate may have particular depth in statistics, software, visualization, experimentation, or a domain; no single interview should require equal mastery of every area. Match prompts to the role, ask candidates to state assumptions, and invite them to explain how they would check their conclusions. The source offers questions, not a validated test or scoring threshold.
For each relevant prompt, listen for a clear explanation, a practical example, trade-offs, and awareness of failure modes. A definition alone may show familiarity with a term, but it does not demonstrate how someone applies it.
Modeling, validation, and error
1. How would you validate a multiple-regression model for a quantitative outcome?
A useful answer should clarify the prediction goal and data structure before choosing a validation method. Ask how the candidate would keep training and evaluation data separate, select metrics appropriate to the use case, inspect residuals and assumptions, and check whether the evaluation reflects the way the model will be used. Strong reasoning also addresses leakage, overfitting, and what would change if observations were grouped or time-ordered.
#1 Best Overall
2. What is overfitting, and how might you reduce its risk?
Overfitting occurs when a model captures patterns in its training data that do not generalize. The companion article describes it as finding chance results that cannot be reproduced in subsequent studies. Ask the candidate to distinguish training performance from performance on genuinely unseen data, then explain relevant safeguards. The companion discusses regularization, randomization testing, nested cross-validation, false-discovery-rate adjustment, and a reusable holdout. Which methods matter depends on the analysis; merely naming them is not a substitute for explaining their purpose.
3. What are precision and recall, and when might you prioritize one?
Look for an explanation tied to the consequences of different errors. Ask what a false positive and a false negative would mean in the application, and how that affects the preferred balance. A strong candidate should connect the choice to the task rather than claim one measure is always best.
4. How do false positives and false negatives differ?
Ask the candidate to define each in a concrete scenario and explain which error is more costly. This tests whether they can translate a model’s classification outcome into the decisions people actually face.
Statistical reasoning and study design
5. What is statistical power?
Ask the candidate to explain the idea in the context of a study they might run, including what the study is trying to detect and what assumptions shape the design. Follow up with how they would respond if the available sample or practical constraints limit what can be detected. The source provides this as a topic prompt, not a prescribed answer or scoring rule.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →6. What is selection bias, why does it matter, and how can you avoid it?
A useful answer identifies how the observed sample could differ from the population or process the analysis is meant to describe. Ask for an example, the direction in which conclusions might be distorted, and how the candidate would detect or reduce the problem. The best discussion makes the limits of any correction explicit.
7. How would you design an experiment to answer a question about user behavior?
The companion article illustrates the question with page load time and user satisfaction. A candidate could identify load time as a factor to vary, satisfaction or a defined behavioral measure as an outcome, and page variants to compare. Ask how they would make the comparison informative, what else could affect the outcome, and how they would interpret the result. The example is illustrative, not a universal experimental protocol.
Rank #3
8. How can repeated hypothesis testing mislead?
When many hypotheses are tested without suitable statistical control, some apparent findings may arise by chance and fail to hold up on repetition. Ask what safeguards fit the analysis. The companion names approaches including simple hypotheses, regularization, randomization testing, nested cross-validation, false-discovery-rate adjustment, and a reusable holdout. A sound answer should explain the risk each approach addresses rather than recite a list.
Data structure and modeling choices
9. What is the difference between long and wide data?
Ask the candidate to describe how observations and features are arranged and to sketch a small example if that helps. The companion contrasts “tall” data—with many more records than features—with “wide” data, which has relatively few records and many features. It warns that methods suited to tall settings can overfit in wide ones. Follow up by asking how the data shape changes the modeling strategy; feature reduction, including Lasso, is one topic the companion raises.
Free tools Windows power users keep installed
One-click scans. No signup required.
10. How would you handle outliers or rare events?
Ask what makes a value an outlier in the specific context, whether it could be a measurement or data-quality issue, and what changes if it is a legitimate rare event. A careful answer avoids automatic deletion and explains how the choice affects analysis and conclusions.
Rank #4
11. How would you approach a recommendation-system problem?
Ask the candidate to clarify the recommendation goal, the available data, and how success would be assessed. Strong reasoning should make assumptions visible and distinguish an offline measure from whether recommendations are useful to people in practice.
Interpreting evidence and communicating results
12. How would you interpret a statistic reported in a published study?
Ask what the statistic measures, which population and conditions it describes, and what conclusions it does—and does not—support. A capable candidate should distinguish an observed association or estimate from a broader causal claim unless the study design justifies that claim.
13. How would you choose a visualization for a data question?
Ask what comparison or pattern the audience needs to see, who the audience is, and how the chart could mislead. A strong answer links chart choice to the data and communication goal, rather than treating visualization as decoration.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Best Value
Turn the prompts into a role-specific interview
Fogg’s 20-question article samples multiple areas; it is not necessary to ask every question in every interview. Choose questions that reflect the actual work and vary the depth to suit the role.
- For applied modeling: emphasize validation, overfitting, error trade-offs, and data structure.
- For experimentation or product analysis: emphasize power, selection bias, experimental design, and interpreting results.
- For communication-focused work: emphasize evidence interpretation and visualization, alongside the technical skills the role requires.
- For any role: ask for a work example, the assumptions involved, what could go wrong, and how the candidate would know whether the result held up.
These are practical ways to adapt the prompts, not a comparative framework validated by a hiring study. A candidate’s answer should be judged against the job’s requirements and demonstrated reasoning, not against an unsupported universal definition of a “real” data scientist.
Source and further reading
Andrew Fogg’s “20 Questions to Detect Fake Data Scientists” was published by KDnuggets on January 1, 2016. Gregory Piatetsky’s companion article, “Answers to 20 Questions to Detect Fake Data Scientists,” discusses selected answers and technical context. It also points to Statistical Learning with Sparsity: The Lasso and Generalizations as further technical background on feature reduction.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

