Scenarios 16–20
16 — Rare fraud records
Situation: 0.2% extreme transactions; team wants to delete them before fraud modeling.
Primary: outlier interpretation tied to purpose.
Best: investigate; they may be valid rare events and the target signal.
Weaker: rarity ≠ error.
Changed: investigation proves sensor/data corruption → DQ correction/removal becomes appropriate.
Source: pp. 488–489.
17 — Pretty but misleading chart
Situation: truncated axis and missing comparison group exaggerate a small effect.
Primary: visualization validity/neutral communication.
Supporting: Governance, Ethics.
Best: statistically valid, neutral presentation with necessary context/assumptions.
Weaker: prettier graphics do not fix biased framing.
Changed: axis/context accurately convey effect and audience understands scale → misleading-presentation problem removed.
Source: pp. 489–490, 497–498.
18 — Millisecond price tag
Situation: accurate failure model; millisecond architecture costs millions; maintenance decisions occur every four hours.
Primary: feasibility/cost vs business latency need.
Best: choose architecture consistent with actionable timing; gate deployment on value/feasibility.
Weaker: fastest system confuses capability with requirement.
Changed: failure can cause immediate catastrophic harm and action must be milliseconds → low latency may be justified.
Source: pp. 484–485, 489–491, 496–497.
19 — MPP or file landing
Situation: petabytes of mixed raw logs need cheap flexible landing; complex low-latency SQL is not immediate need.
Primary: architecture fit for landing/storage.
Best: distributed file-based environment; add analytical engines as requirements demand.
Weaker: MPP selected solely for structured-query performance may not fit cost/landing need.
Changed: repeated high-performance parallel structured analytics becomes dominant → MPP more attractive.
Source: pp. 491–494.
20 — DQ deferred
Situation: integrate five providers, build model, assess quality afterward.
Primary: DQ positioned too late.
Supporting: DQ, Integration, Metadata.
Best: profile/classify/assess each source before integration; use findings for mapping, selection, feasibility.
Weaker: late DQ embeds uncertainty and hides root causes.
Changed: all providers already governed/profiled and meet known quality thresholds → integration/alignment becomes next.
Source: pp. 486–488, 499–500.