Diagnostics C14-025–036
C14-025 — A
Why correct: source context should be preserved while acquiring/ingesting, including origin, size, currency, structure/content, classification/profiling, and quality context.
A: Correct. B: filename alone is inadequate. C: not all sources have a target label. D: visualization comes later.
Source: pp. 486–487, 499 · Confusion: Ingest vs Metadata-after-the-fact · Tag: sequence/process
C14-026 — B
Why correct: source DQ findings guide whether data is usable and how it should be mapped/aligned before bad evidence is embedded downstream.
A: Integrated data is not automatically unprofileable. B: Correct. C: DQ and Metadata complement one another. D: fitness, not perfect accuracy, is the standard.
Source: pp. 486–488, 499–500 · Confusion: DQ before vs after integration · Tag: sequence/process
C14-027 — C
Why correct: a matching customer key does not align observations that represent different time grains.
A: ETL/ELT order does not fix time semantics. B: more rows do not align grain. C: Correct. D: learning method is unrelated.
Source: pp. 487–488 · Confusion: Key match vs analytical alignment · Tag: missed qualifier
C14-028 — D
Why correct: trusted identifiers/entity resolution and Master/Reference Data context help unify identity across cookie/customer/account IDs.
A: visualization is downstream. B: not a sentiment problem. C: scale does not establish identity. D: Correct.
Source: pp. 487–488 · Confusion: Integration vs MDM/RDM support · Tag: cross-domain confusion
C14-029 — A
Why correct: Metadata establishes meaning/provenance/context; DQ establishes fitness/trust of values for the use.
A: Correct. B: reverses roles. C: both are much broader. D: they are distinct disciplines.
Source: pp. 499–500 · Confusion: Metadata vs Data Quality · Tag: confusion pair
C14-030 — B
Why correct: the discovered failures should have been exposed through source profiling/DQ and semantic alignment before integration/modeling.
A: multiple sources are explicitly supported. B: Correct. C: analytics category cannot fix source evidence. D: serving speed is irrelevant.
Source: pp. 485–489, 499–500 · Confusion: DQ/alignment sequence · Tag: sequence/process
C14-031 — C
Why correct: over-fitting means learning training-specific noise/patterns too closely and failing to generalize.
A: slowness is feasibility/performance. B: small data alone does not define over-fit. C: Correct. D: visualization is unrelated.
Source: pp. 488–490, 494–495 · Confusion: Over-fitting vs performance · Tag: vocabulary confusion
C14-032 — D
Why correct: after choices are made, an untouched test set should provide the final independent generalization estimate.
A: training shaped the model. B: validation shaped selection/tuning. C: Metadata is not an evaluation partition. D: Correct.
Source: pp. 489–490, 494–495 · Confusion: Training vs Validation vs Test · Tag: confusion pair
C14-033 — A
Why correct: once test results influence feature/model choices, the test set has joined selection and is no longer independent.
A: Correct. B: direct parameter fitting is not required for selection leakage. C: learning type is unrelated. D: visualization cannot restore independence.
Source: pp. 489–490, 494–495 · Confusion: Validation vs Test · Tag: changed-answer behavior
C14-034 — B
Why correct: outliers can be valid rare events and may be exactly the fraud signal the model is intended to find.
A: rarity does not prove invalidity. B: Correct. C: averaging can destroy signal. D: architecture does not determine validity.
Source: pp. 488–489 · Confusion: Outlier vs Error · Tag: best-action error
C14-035 — C
Why correct: visualization must fit a defined question/audience and communicate neutrally with enough context for valid interpretation.
A: complexity ≠ fitness. B: explanatory context can be necessary. C: Correct. D: maximizing variables can obscure meaning.
Source: pp. 489–490, 497–498 · Confusion: Visualization sophistication vs fitness · Tag: overthinking
C14-036 — D
Why correct: the source calculation can be correct while the visual framing misleads; this is interpretation/visualization governance failure.
A: correct underlying values mean this is not automatically DQ. B: “chart” does not make the main issue Metadata. C: Volume is irrelevant. D: Correct.
Source: pp. 497–498 · Confusion: Correct calculation vs misleading communication · Tag: cross-domain confusion