Skip to content
WeavidenceInsights
Follow

Concept ExplainerScientific Learning

What counts as evidence of demonstrated competence?

In brief: Evidence of competence is a chain, not a checkmark. It should distinguish knowledge from application and performance, sample the capability across appropriate tasks and contexts, use transparent criteria and multiple observations where the claim is consequential, and show that the learner can transfer reasoning beyond rehearsed examples.

Question

What actually counts as evidence that someone is competent, rather than merely exposed to the material?

Answer in brief

Competence is not demonstrated by attendance, content completion, or one high score. Those may be valid evidence of participation or knowledge, but the claim "this person can perform" is broader. A defensible competence claim normally requires evidence that the learner can select and apply knowledge in an appropriate task, transfer reasoning beyond rehearsed examples, perform against explicit criteria, and do so with enough consistency across observations and contexts for the intended decision.

Miller's influential framework distinguishes knows, knows how, shows how, and does.[1] It is a useful map of increasingly performance-proximal evidence, not a rule that one task at a higher level proves competence. Modern programmatic-assessment approaches therefore combine multiple low- and high-stakes observations into a coherent judgment rather than treating one assessment event as a verdict.[2][3]

Key points

  • Completion and competence answer different questions. Completion shows that a learner reached the end of an activity; competence concerns what the learner can select, integrate, and perform.
  • Assessment method must match the claim. Recall questions can assess knowledge. They cannot, by themselves, establish performance in an unfamiliar or real setting.
  • Transfer is an empirical property. Practice and testing can improve transfer, but the effect varies with the distance between training and target tasks, the cues available, and how knowledge was learned.[4]
  • One observation is rarely enough for a consequential judgment. Cases, assessors, settings, and occasions sample different parts of performance.
  • Evidence should be accumulated and interpreted. Programmatic assessment treats individual observations as data points within a larger decision process, with feedback and proportional stakes.[2]
  • Authenticity is not the same as validity. A realistic simulation can still measure the wrong construct, use weak criteria, or sample too narrow a task.
  • The decision determines the required evidence. A formative claim that a learner is ready for more practice requires less evidence than certification for independent professional responsibility.

Scope and exclusions

This article explains the evidentiary logic of competence assessment in scientific, methodological, and professional learning. It does not provide a credentialing standard for a specific profession, replace psychometric or legal requirements, define workplace authorization, or determine whether an individual is competent. Miller's model and much of the cited assessment literature originated in health-professions education; applying the principles to broader scientific learning requires attention to the target capability and context.

Four different claims about learning

Knows

The learner can recall or recognize relevant facts, concepts, terminology, and rules. Selected-response and short-answer questions can provide efficient samples of this level when they are well designed.

Evidence at this level supports a knowledge claim. It does not show that the learner can identify when the knowledge is relevant or use it under uncertainty.

Knows how

The learner can explain, choose, or justify an approach in a described problem. A strong task presents a situation that requires selecting and combining knowledge rather than reproducing a memorized phrase.

This is where transfer begins to matter. The surface details should differ from practice enough to require recognition of the underlying principle, while remaining within the capability the programme intended to teach.

Shows how

The learner performs in an observed, controlled setting: a simulation, structured practical task, analysis exercise, oral defence, or supervised review. The assessment can evaluate process, reasoning, communication, and output against explicit criteria.

A demonstration is closer to performance, but it is still a sample under constructed conditions. Familiarity with the scenario, coaching, assessor variation, and simulation constraints all affect interpretation.

Does

Evidence comes from performance in real practice. This may include work products, direct observation, audit trails, decisions, outcomes, and structured feedback from people who observed the work.

Real-world evidence is highly relevant but not automatically clean: cases vary, opportunities are uneven, supervision differs, outcomes are influenced by teams and systems, and unsafe performance cannot be permitted merely to create an assessment opportunity.

Miller's pyramid makes these claims visibly different.[1] It should not be read as saying that every capability must be assessed only at the top, or that one "does" observation settles the question.

Why a single pass is weak evidence

Performance varies across tasks. A learner may reason well about confounding and poorly about measurement error; may produce a correct model with clean data and fail when variables are ambiguously coded; or may perform well when prompted but omit the same safeguard independently.

A single assessment result combines at least three things:

observed result
= capability relevant to the task
+ task and context effects
+ measurement error

Increasing difficulty does not solve this sampling problem. One extremely hard case is still one case. The stronger approach is to collect multiple observations that differ deliberately in relevant ways and then interpret the pattern.

Programmatic assessment formalizes this idea: individual observations are optimized for learning and information, while higher-stakes decisions draw on a rich body of evidence rather than one examination event.[2] The quality of the programme depends on the coherence of the evidence, feedback, aggregation, and decision process — not merely on the number of data points.

Transfer: the capability behind an unfamiliar problem

Transfer occurs when learning influences performance in a task or context that is not identical to the learning event. It is often invoked too casually. An assessment does not test meaningful transfer merely because the names or numbers changed.

The distance between learning and assessment can vary:

  • near transfer: the same principle in a closely related case;
  • moderate transfer: different surface features requiring the learner to identify the same underlying structure;
  • far transfer: applying principles across substantially different domains, contexts, or representations.

A meta-analysis of test-enhanced learning found that prior testing can improve performance on transfer tasks, but effects depend on moderators such as the relation between practice and transfer material.[4] This supports using retrieval and application practice, while warning against assuming that success on practiced items automatically generalizes.

To test transfer honestly:

  1. define what should remain invariant across contexts;
  2. vary features that should not determine the answer;
  3. avoid giving the learner the exact cue used during teaching;
  4. require explanation or an inspectable product, not only a final choice;
  5. include more than one transfer task;
  6. state how far beyond the trained context the claim is intended to extend.

A competence evidence chain

A defensible system can represent competence as a chain:

intended capability
→ assessment blueprint
→ tasks and contexts
→ observed performance
→ criteria and judgments
→ aggregation across observations
→ decision with stated scope and uncertainty

Intended capability

Write the capability as an observable integration of knowledge, reasoning, and action. "Understand study design" is too vague. "Given a public-health question, compare plausible study designs, identify material trade-offs, and justify a selection without overstating causal support" is assessable.

Assessment blueprint

Map the capability to a representative set of tasks. The blueprint should cover important variations, not simply the easiest content to score automatically.

Tasks and contexts

Use a deliberate mixture of knowledge, application, demonstration, and — where appropriate and ethical — workplace evidence. The mix should reflect the claim and stakes.

Observed performance

Retain the actual response, reasoning, work product, or structured observation. A binary pass flag discards information needed for feedback and later review.

Criteria and judgments

Criteria should identify what good performance requires and where professional judgment remains necessary. Assessor training, calibration, and narrative feedback can matter as much as a numerical rubric.

Aggregation

Combine evidence transparently. Averages can hide a critical recurring failure; one excellent task can mask missing coverage. Decision rules should consider patterns, recency, critical dimensions, and contradictory evidence.

Decision and scope

State exactly what has been concluded. "Demonstrated this capability across four synthetic study-review tasks under supervised conditions" is more honest and useful than a universal label of "competent."

Worked example: evaluating an observational study

A programme teaches learners to evaluate causal claims made from routine health data.

Knowledge evidence

  • define confounding, selection bias, and measurement error;
  • recognize common study designs;
  • identify the meaning of an adjusted estimate.

This supports a knows claim.

Application evidence

The learner receives several unfamiliar short scenarios and must:

  • identify the target causal question;
  • propose plausible confounders;
  • distinguish confounding from selection and measurement problems;
  • explain why adjustment for an available variable may or may not address the problem.

This supports a knows how claim if the tasks genuinely require transfer.

Demonstration evidence

The learner reviews a complete synthetic study package containing a protocol, data dictionary, analysis plan, results, and missing-data report. They produce a structured critique and defend it under questioning. Assessors evaluate both the reasoning process and the final conclusions.

This supports a shows how claim for the sampled conditions.

Practice evidence

In a real role, the learner's reviews are sampled over time and compared with independent expert review, revision requests, and the quality of the final record. Confidentiality, supervision, and patient or organizational safety take priority over assessment convenience.

This may support a does claim, but only within the tasks, context, oversight, and period actually observed.

Aggregated decision

A defensible conclusion might be:

The learner consistently identifies major threats to causal interpretation and justifies study-design critiques across six unfamiliar synthetic cases and two supervised real reviews. Evidence is insufficient for independent review of complex longitudinal causal models.

That conclusion is narrower than "competent in causal inference" and therefore more informative.

Evidence and method

Miller's framework supplies the distinction between knowledge, application, demonstration, and practice.[1] Van der Vleuten and Schuwirth argue that professional competence is complex, context-dependent, and best approached through complementary assessment methods rather than a search for one perfect instrument.[3] The programmatic-assessment model extends that principle into an assessment system in which individual observations contribute to learning and later high-stakes decisions.[2] The transfer discussion is grounded in the meta-analysis by Pan and Rickard, which synthesizes how test-enhanced learning generalizes beyond practiced material.[4]

These sources do not define one universal competence standard. They support the architecture of an evidence argument; professions and institutions must still define capabilities, acceptable performance, stakes, and governance.

Practical implications

When a course, credential, or platform claims to demonstrate competence, ask:

  • What capability is being claimed?
  • At which level — knowledge, application, demonstration, or practice — was it assessed?
  • How many tasks, settings, occasions, and assessors contributed?
  • Were tasks sampled from a blueprint?
  • Was transfer tested or only rehearsed material repeated?
  • Is the actual performance record retained?
  • How were conflicting observations handled?
  • Who made the final judgment and under what standard?
  • What is the scope and expiry of the conclusion?
  • What evidence would trigger remediation or reassessment?

A completion certificate can be completely honest when it is labelled as completion. The problem begins when evidence of exposure is silently promoted into evidence of capability.

Limitations and uncertainty

Most cited models come from health-professions education. Their application to scientific and methodological learning is conceptually relevant but must be adapted rather than copied. Competence is also socially and institutionally defined: acceptable performance depends on role, risk, supervision, resources, and regulation. More observations do not guarantee validity when all tasks are narrow, assessors share the same bias, or the intended capability is poorly defined. Real-world outcomes are valuable but may be too delayed or confounded to attribute to one learner. No assessment architecture removes the need for professional judgment and periodic review.

Connection to Weavidence

Weavidence Academy is designed to keep viewed, practised, passed, and demonstrated as distinct states and to retain the evidence behind them. That separation can prevent completion from being overstated as competence. The product remains at foundation stage; no claim is made that its assessment model has yet demonstrated validity, reliability, educational effectiveness, or fitness for credentialing in practice.

References

  1. [1] The assessment of clinical skills/competence/performance Miller GE. 1990. Academic Medicine, 65(9 Suppl), S63–S67 Source · PMID 2400509
  2. [2] A model for programmatic assessment fit for purpose van der Vleuten CPM, Schuwirth LWT, Driessen EW, Govaerts MJB, Heeneman S. 2012. Medical Teacher, 34(3), 205–214 Source · DOI · PMID 22364452
  3. [3] The assessment of professional competence: building blocks for theory development van der Vleuten CPM, Schuwirth LWT. 2010. Best Practice & Research Clinical Obstetrics & Gynaecology, 24(6), 703–719 Source · DOI · PMID 20510653
  4. [4] Transfer of test-enhanced learning: Meta-analytic review and synthesis Pan SC, Rickard TC. 2018. Psychological Bulletin, 144(7), 710–756 Source · DOI · PMID 29733621

Revision history

No revisions since original publication on .