Diagnostics C14-037–048
C14-037 — A
Why correct: MPP shared-nothing partitions data/computation across nodes with dedicated resources for scalable parallel analytics.
A: Correct. B: shared disk/memory contradicts shared-nothing. C: removes distributed parallel design. D: streaming is not defining.
Source: pp. 491–493 · Confusion: MPP vs shared-resource system · Tag: vocabulary confusion
C14-038 — B
Why correct: flexible distributed file environments fit very large varied-data landing/storage where immediate interrogative analytics is secondary.
A: MPP is strong for parallel analytics, not automatically best landing choice. B: Correct. C: cannot fit scale. D: serving is not storage substrate.
Source: pp. 493–494 · Confusion: MPP vs distributed file-based · Tag: architecture selection
C14-039 — C
Why correct: the conceptual distributed workflow is Map → Shuffle → Reduce.
A: ETL. B: evaluation partitions. C: Correct. D: service layers.
Source: pp. 493–494 · Confusion: MapReduce vs ETL/layers · Tag: vocabulary confusion
C14-040 — D
Why correct: in-database algorithms run analysis near stored data, reducing movement and using platform processing resources.
A: does not remove DQ/Metadata. B: cannot guarantee generalization. C: does not automatically relationalize all data. D: Correct.
Source: pp. 493–494 · Confusion: In-database vs data movement · Tag: vocabulary confusion
C14-041 — A
Why correct: repeated high-performance parallel queries over partitionable structured data match MPP's highlighted strength.
A: Correct. B: file landing can coexist but is weaker for the stated analytic workload. C: consumes results. D: terminology management does not execute queries.
Source: pp. 491–493 · Confusion: MPP vs distributed file-based · Tag: architecture selection
C14-042 — B
Why correct: leased/cloud capacity can support exploration and economic/technical feasibility before permanent commitment.
A: strategy still leads technology. B: Correct. C: temporary does not remove governance. D: test independence still matters.
Source: pp. 493–494, 496–497 · Confusion: Cloud exploration vs tool-first strategy · Tag: best-action error
C14-043 — C
Why correct: source variation makes disciplined Metadata more important so users know what exists, origin, meaning, and value/context.
A: Chapter explicitly emphasizes Metadata. B: context is needed at ingest and throughout. C: Correct. D: Metadata is broader than size/location.
Source: pp. 472–473, 499 · Confusion: Metadata principle · Tag: vocabulary confusion
C14-044 — D
Why correct: high-velocity data can still be wrong, sensitive, biased, misinterpreted, or badly sourced; fast propagation can amplify failure.
A: governance does not guarantee speed. B: velocity is explicitly a V. C: real-time predictive use exists. D: Correct.
Source: pp. 496–500 · Confusion: Velocity vs controls · Tag: overgeneralization
C14-045 — A
Why correct: uncontrolled combination of internal customer and external location data is a sourcing/sharing/access/security/privacy governance problem with recombination risk.
A: Correct. B: performance tuning does not govern access/privacy. C: analytics category is secondary. D: sample size is irrelevant.
Source: pp. 497–499 · Confusion: Governance vs analytics type · Tag: cross-domain confusion
C14-046 — B
Why correct: a good pilot is only part of readiness; staff, procurement, ownership, business adoption, and operational/economic feasibility must exist.
A: accuracy is already described as good. B: Correct. C: Metadata is not the blocking fact. D: descriptive maturity is unrelated.
Source: pp. 496–497 · Confusion: Prototype success vs organizational readiness · Tag: best-action error
C14-047 — C
Why correct: revenue, cost reduction, avoided threat, useful models/patterns, and new initiatives connect work to business value rather than activity.
A: loading activity only. B: capacity health only. C: Correct. D: query activity does not by itself prove value.
Source: pp. 500–501 · Confusion: Activity metrics vs value metrics · Tag: confusion pair
C14-048 — D
Why correct: deploy-and-monitor is broader than prediction accuracy: source/Metadata drift, operational effectiveness, adoption, business value, and refinement matter.
A: stable accuracy does not prove current meaning/usefulness. B: uptime is necessary but insufficient. C: deployment does not end the lifecycle. D: Correct.
Source: pp. 490–491, 499–501 · Confusion: Accuracy vs ongoing value/Metadata drift · Tag: monitoring gap