Skip to content

Diagnostics C14-013–024

C14-013 — A

Why correct: labeled spam/not-spam outcomes provide the known target for supervised learning.
A: Correct. B: Unsupervised has no known target. C: Reinforcement uses reward/feedback. D: Association finds co-occurrence.
Source: pp. 480–481 · Confusion: Supervised vs Unsupervised · Tag: vocabulary confusion

C14-014 — B

Why correct: no authoritative segment labels exist; the task is hidden-pattern/group discovery.
A: Needs known target. B: Correct. C: No action/reward loop. D: No recommended action.
Source: pp. 480–482 · Confusion: Unsupervised vs Supervised · Tag: confusion pair

C14-015 — C

Why correct: association relates/co-locates elements; clustering groups similar observations.
A: Not defined by labels. B: Not future-vs-past. C: Correct. D: Both are analytic/mining techniques.
Source: pp. 481–482 · Confusion: Association vs Clustering · Tag: confusion pair

C14-016 — D

Why correct: sentiment depends on linguistic context, negation, word combinations, and other contextual meaning.
A: Sentiment targets text. B: Speed is not the deciding weakness. C: Not MDM. D: Correct.
Source: pp. 482–483 · Confusion: Sentiment vs keyword frequency · Tag: vocabulary confusion

C14-017 — A

Why correct: “products commonly appear together” is a co-occurrence/association question.
A: Correct. B: Clustering groups similar cases. C: No action recommendation. D: Reduction simplifies variables/data.
Source: pp. 481–482 · Confusion: Association vs Clustering · Tag: technique selection

C14-018 — B

Why correct: the defining clue is iterative reward/feedback toward a goal.
A: No labeled target examples. B: Correct. C: Not merely hidden-group discovery. D: Not historical summary.
Source: pp. 480–481 · Confusion: Reinforcement vs Supervised/Unsupervised · Tag: missed qualifier

C14-019 — C

Why correct: Chapter 14 begins with strategy/business need so source, platform, model, and architecture choices have a purpose.
A: Tool-first reasoning. B: Source acquisition follows the problem. C: Correct. D: Model performance does not define the business need.
Source: pp. 484–485, 496–497 · Confusion: Strategy vs tool selection · Tag: sequence/process

C14-020 — D

Why correct: granularity is the level of detail represented by the data.
A: Velocity=speed. B: Veracity=trust. C: Volatility=change/useful life. D: Correct.
Source: pp. 485–487 · Confusion: Granularity vs V characteristics · Tag: vocabulary confusion

C14-021 — A

Why correct: uncertain sampling and changing definitions create reliability, semantic, and population-bias risks that must be assessed before relying on the feed.
A: Correct. B: Ingest-first postpones the source decision. C: Speed does not make unreliable evidence acceptable. D: field names are insufficient context.
Source: pp. 485–487 · Confusion: Source volume/cost vs reliability · Tag: best-action error

C14-022 — B

Why correct: training only on app users may not represent all customers, so the model may be biased for the deployment population.
A: Size is not the key fact. B: Correct. C: No swamp symptoms. D: Serving is not the primary issue.
Source: pp. 486, 489, 498 · Confusion: Bias vs Data Volume · Tag: missed qualifier

C14-023 — C

Why correct: recombination can reveal identity/sensitivity that was not visible in either source alone.
A: Join accuracy may matter but does not address the new privacy fact. B: Provenance is not the cause. C: Correct. D: Row count is irrelevant.
Source: pp. 486, 498–499 · Confusion: Privacy vs Metadata/DQ · Tag: cross-domain confusion

C14-024 — D

Why correct: source values can be reliable yet unusable/governance-risky when origin, meaning, lineage, and intended use are unknown.
A: Stem says quality is strong. B: No prescription issue. C: Architecture cannot create missing context. D: Correct.
Source: pp. 472–473, 486–487, 499 · Confusion: Metadata vs Data Quality · Tag: cross-domain confusion

← Diagnostics 001–012 · Diagnostics 025–036 →