Skip to content

Rapid Recall 41–60

  1. What is recombination risk?
  2. What should happen before integrating a new source?
  3. Why is Data Quality not a last-step cleanup activity?
  4. What is the role of Master/Reference Data in alignment?
  5. Hypothesis quality vs input-data quality?
  6. Why pre-populate a predictive model with history?
  7. What does model training mean?
  8. What is over-fitting?
  9. Training vs validation set?
  10. Validation vs test set?
  11. Why does repeated test-set reuse weaken evaluation?
  12. What does K-fold/cross-validation do?
  13. Why should outliers be investigated before deletion?
  14. What makes a visualization fit for purpose?
  15. Static vs interactive visualization?
  16. When should a model be deployed?
  17. What must be monitored after deployment?
  18. MPP shared-nothing in one sentence.
  19. Distributed file-based architecture in one sentence.
  20. MapReduce: name the three conceptual phases.

← Recall 21–40 · Key 01–20 →