Article 3: What Counts as Evidence in an Operations Diagnostic?

Use evidence and triangulation to evaluate those explanations.

Domain 1. Article 3.

By Edwin Angulo | Co-Founder, Angulo & Morsa Legacy Consulting LLC

Published July 27, 2026 | Last reviewed July 27, 2026

EXECUTIVE TAKEAWAY

Evidence in an operations diagnostic is information that helps support, weaken, or overturn a specific explanation for a defined performance gap.

Direct system data, transaction records, observation, time studies, customer behavior, employee interviews, and management interpretation can all contribute useful evidence. They do not provide the same type or strength of information, and none should automatically be accepted as complete or reliable.

A single source rarely proves root cause.

A credible diagnosis usually requires triangulation: comparing several independent or complementary sources to determine whether they describe the same condition, reveal different parts of the process, or contradict one another in ways that require further investigation.

Article 2 established that operational diagnosis should begin with a measurable problem statement:

  • What is happening now?

  • What should be happening?

  • How large is the gap?

  • Which process and population are affected?

  • Why does the gap matter?

  • What is included and excluded from the investigation?

Article 1 explained the next step: develop competing, testable explanations for that gap instead of accepting management’s first theory as fact.

This article addresses the next requirement:

What information would actually count as evidence for or against those explanations?

The sequence is:

  1. Define the measurable operational gap.

  2. Develop competing hypotheses that could explain the gap.

  3. Identify the evidence needed to evaluate each hypothesis.

  4. assess the quality and limitations of that evidence.

  5. Compare findings across multiple sources.

  6. Determine which explanations remain credible.

  7. Test the proposed countermeasure.

  8. Measure whether the original gap improves.

Without a defined problem, data collection becomes unfocused.

Without competing hypotheses, evidence collection becomes biased toward the preferred explanation.

Without reliable evidence, root-cause analysis becomes organized speculation.

The operating principle is simple:

Evidence should be collected to test a defined claim, not merely to support what management already believes.

A dashboard may show that performance declined.

An employee may explain why they believe it declined.

A manager may identify the person or system believed to be responsible.

A customer complaint may illustrate the consequences.

A time study may reveal where work is waiting.

Each source may contribute something important.

None necessarily tells the whole story.

Data and Evidence Are Not the Same Thing

Businesses produce large amounts of data:

  • CRM records

  • Scheduling reports

  • Financial statements

  • Call logs

  • Emails

  • Work orders

  • Customer reviews

  • Employee surveys

  • Productivity reports

  • Website analytics

  • Spreadsheets

  • Meeting notes

The existence of data does not make it relevant evidence.

Data becomes evidence only when it is connected to a defined question or hypothesis.

Suppose the problem statement is:

During the 90 days ending June 30, 2026, 69% of qualified inquiries received documented follow-up within one business day, compared with the company’s target of 95%.

Management suspects that sales employees are ignoring CRM notifications.

A report showing total monthly website traffic is data, but it does not directly test whether employees ignored notifications.

A CRM audit trail showing when inquiries were received, assigned, viewed, contacted, and closed may be relevant evidence.

Employee interviews about how notifications are received and prioritized may provide explanatory evidence.

Direct observation of the intake and assignment workflow may reveal that inquiries do not reach employees until several hours after submission.

The evidence must bear on the claim being tested.

EVIDENCE PRINCIPLE

The amount of information collected does not determine the strength of a diagnosis.

The relevant question is whether the information is sufficiently reliable, direct, representative, and connected to the decision.

A business can possess millions of records and still lack the evidence needed to explain a specific operational problem.

Evidence Does Not Always Mean Proof

In ordinary business conversation, leaders may use the word “proof” loosely:

  • The dashboard proves employees are not following the process.

  • The complaint proves customers are unhappy.

  • The interview proves the software is the problem.

  • The sales decline proves the new pricing failed.

Most operational evidence is not that conclusive.

Evidence changes the degree of confidence in an explanation.

It may:

  • Support an explanation

  • Weaken an explanation

  • Contradict an explanation

  • Reveal that the explanation is incomplete

  • Identify a more plausible competing explanation

  • Show that the available information is insufficient

  • Expose a problem in the measurement system itself

A system report may establish that an action was not documented.

It may not establish that the action never happened.

An employee interview may establish that a workaround exists.

It may not establish how frequently the workaround occurs.

Customer abandonment may establish that customers exited the process.

It may not establish why they exited.

A strong diagnostic distinguishes what each source can establish from what the source cannot establish on its own.

A Practical Evidence Hierarchy

The following hierarchy provides a general starting point for evaluating operational evidence:

  1. Direct system data

  2. Transaction or process records

  3. Direct observation

  4. Time studies

  5. Customer behavior

  6. Employee interviews

  7. Management interpretation

  8. Anecdotes and assumptions

This is not a rigid universal ranking.

A higher category is not automatically more reliable than every source beneath it.

A CRM database with missing records, inconsistent definitions, overwritten timestamps, and weak access controls may be less trustworthy than a carefully conducted observation or review of original transaction records.

An employee interview may reveal a hidden workaround that no formal system captures.

Customer behavior may provide the clearest evidence of an outcome even when internal records appear complete.

The hierarchy reflects a general preference for evidence that is:

  • Direct rather than indirect

  • Recorded close to the event

  • Independently verifiable

  • Consistently defined

  • Less dependent on memory

  • Less vulnerable to interpretation

  • Available across enough transactions to identify patterns

The strength of the evidence still depends on its quality and relevance to the question.

1. Direct System Data

Direct system data includes information generated or stored by the systems through which work is performed.

Examples include:

  • CRM event histories

  • ERP transaction logs

  • Scheduling timestamps

  • Call-system records

  • Website-form submissions

  • Workflow audit trails

  • Queue histories

  • User-access records

  • Status changes

  • Assignment times

  • Completion times

  • Inventory movements

  • Billing transactions

  • Automated system alerts

  • Customer-support ticket histories

Direct system data is often the strongest starting point because it can show activity across a large number of transactions.

It may allow the diagnostic team to determine:

  • When an event occurred

  • Who or what system handled it

  • How long the transaction remained in each status

  • Whether required fields were completed

  • Which channels, locations, employees, or transaction types differ

  • Whether the problem is consistent or concentrated

  • Whether performance changed over time

  • Where work accumulated or disappeared

For the inquiry-follow-up problem, direct system data might show:

  • The time each inquiry was received

  • The time it was qualified

  • The time it was assigned

  • The employee assigned

  • The time of first documented contact

  • The communication channel

  • The final status

  • Whether the record was reassigned

  • Whether the record was reopened

  • Whether the customer converted

This can help reconstruct the digital history of the process.

But system data has an important limitation:

A system records what it was designed and used to record—not necessarily everything that actually happened.

The data may be incomplete or misleading when:

  • Employees bypass the system

  • Required fields are optional

  • Timestamps are overwritten

  • Employees close tasks without completing the intended work

  • Status definitions are inconsistent

  • Records are duplicated

  • Several systems capture different parts of the process

  • Configuration changed during the review period

  • Data were imported incorrectly

  • User access is shared

  • Managers manually modify records

  • The report excludes certain transaction types

  • Employees document work after the fact

A system can produce precise numbers from an unreliable process.

Precision does not guarantee validity.

Before relying on direct system data, the diagnostic should examine:

  • Accuracy

  • Completeness

  • Applicability to the question

  • Data definitions

  • Access controls

  • Required-field rules

  • Missing values

  • Duplicate records

  • Changes in configuration

  • The process used to enter or generate the data

The U.S. Government Accountability Office describes data reliability in terms of accuracy, completeness, and applicability for the intended purpose.

The same principle applies to an operations diagnostic.

The data does not need to be perfect.

The diagnostic team must understand whether it is reliable enough to support the decision being considered.

2. Transaction or Process Records

Transaction and process records are the documents created as work moves through the organization.

Examples include:

  • Work orders

  • Intake forms

  • Customer files

  • Invoices

  • Purchase orders

  • Checklists

  • Approval records

  • Email messages

  • Call notes

  • Calendars

  • Shipping records

  • Service reports

  • Inspection forms

  • Training records

  • Handoff documents

  • Exception reports

  • Complaint records

  • Paper files

  • Signed acknowledgments

These records may confirm, supplement, or challenge the system data.

A system report may indicate that a customer was contacted at 10:14 a.m.

The transaction record may show:

  • An email sent at that time

  • A telephone note

  • A voicemail recording

  • A completed contact form

  • No corresponding evidence of actual contact

Process records are especially useful when the workflow crosses multiple systems or contains manual steps.

They can help reconstruct individual cases and answer questions such as:

  • What information was available at the time?

  • Was the transaction complete?

  • Was the required approval obtained?

  • Were instructions clear?

  • Did the handoff contain all necessary information?

  • Was the same information entered more than once?

  • Did the record move through the expected sequence?

  • Were exceptions documented?

  • Did employees create parallel records outside the primary system?

Transaction records often provide greater detail than a summary dashboard.

They also have limitations.

Records may be:

  • Missing

  • Incomplete

  • Backdated

  • Created after the event

  • Inconsistently formatted

  • Stored in several locations

  • Selected only when something went wrong

  • Written to satisfy documentation requirements rather than describe actual work

  • Copied from one transaction to another

  • Difficult to connect to the correct customer or event

A completed form may prove that a field was entered.

It may not prove that the underlying action occurred correctly.

A checked box does not automatically establish process compliance.

The diagnostic should compare the record with other evidence.

3. Direct Observation

Direct observation examines how the work actually happens.

It may involve:

  • Watching the intake process

  • Following a transaction through several handoffs

  • Observing a scheduling employee

  • Attending a management review

  • Walking through a physical workflow

  • Reviewing how employees use the software

  • Watching how exceptions are handled

  • Observing customer interactions

  • Comparing the documented process with actual behavior

Lean management uses the principle of going to the actual place where work occurs to understand the real condition rather than relying solely on reports or secondhand explanations.

Direct observation helps distinguish the process as designed from the process as performed.

The written procedure may show:

Inquiry received → automatically assigned → employee follows up within one business day.

Observation may reveal:

Inquiry received → shared inbox reviewed several times per day → intake employee copies information into a spreadsheet → manager decides who should receive it → employee receives an email → employee manually creates the CRM record → follow-up begins.

Both descriptions may refer to the same organization.

Only one describes the actual workflow.

A useful metaphor is the difference between a map and the terrain.

The procedure, policy, or flowchart is the map.

The actual work is the terrain.

A map can be accurate, outdated, incomplete, or aspirational.

Direct observation shows where employees actually walk, where they stop, where they turn around, and where they create shortcuts because the official route does not work.

Observation can reveal:

  • Interruptions

  • Waiting

  • Rework

  • Informal approvals

  • Workarounds

  • Duplicate data entry

  • Unclear ownership

  • Missing information

  • Physical constraints

  • Software-navigation problems

  • Unrecorded communication

  • Exception handling

  • Differences among employees

  • Differences between normal and peak periods

Observation also has limitations.

Employees may change their behavior because they know they are being observed.

The observer may see:

  • An unusually calm day

  • An unusually busy day

  • One employee who is not representative

  • A process being performed more carefully than usual

  • Only a small number of transactions

  • Only the visible part of a longer process

Observation should therefore be conducted respectfully and across enough conditions to understand whether the observed behavior is typical.

Its purpose is not to catch employees doing something wrong.

Its purpose is to understand what the operating system requires people to do.

4. Time Studies

A time study measures how long work takes and where time is consumed.

It can distinguish among:

  • Touch time

  • Wait time

  • Queue time

  • Processing time

  • Travel time

  • Rework time

  • Handoff time

  • Approval time

  • Customer-response time

  • System-response time

This distinction matters because elapsed time and working time are not the same.

A customer may wait eight hours for a response even though the transaction requires only twelve minutes of actual employee effort.

The delay may result from:

  • Queue position

  • Batching

  • Prioritization rules

  • Missing information

  • Waiting for approval

  • Workload peaks

  • Unclear assignment

  • System latency

  • Customer availability

  • Service-capacity constraints

A time study can help determine whether a process is slow because the work itself takes a long time or because the work spends most of its time waiting.

For example:

  • Inquiry qualification touch time: four minutes

  • Average wait before qualification: three hours

  • Assignment touch time: one minute

  • Average wait before assignment: five hours

  • First follow-up touch time: eight minutes

The total elapsed time may exceed one business day even though the direct labor requirement is relatively small.

That pattern would point toward queue design, batching, ownership, prioritization, or capacity rather than the complexity of the individual task.

Time studies can also identify:

  • Variation among transaction types

  • Variation among employees

  • Peak-period bottlenecks

  • Unnecessary motion

  • Repeated system navigation

  • Duplicate entry

  • Frequent interruptions

  • Approval delays

  • Tasks that add no customer or control value

Time-study limitations include:

  • Small or unrepresentative samples

  • Observer effects

  • Employees changing pace during measurement

  • Differences in experience

  • Differences in case complexity

  • Learning effects

  • Seasonal demand

  • Unusual staffing

  • Failure to separate normal work from exception work

A time study should not be used as a simplistic tool for forcing every employee to match the fastest observed time.

The purpose is to understand the process and its variation, not to treat people as interchangeable machines.

5. Customer Behavior

Customer behavior includes what customers actually do as they interact with the process.

Examples include:

  • Completing or abandoning an inquiry

  • Responding to follow-up

  • Scheduling an appointment

  • Canceling

  • Failing to appear

  • Purchasing

  • Declining

  • Requesting a refund

  • Returning for another service

  • Submitting a complaint

  • Switching channels

  • Repeatedly contacting the company

  • Escalating to a manager

  • Leaving the process without explanation

Customer behavior can provide stronger evidence than stated preference because it reflects an actual action.

A customer may say that rapid response is important.

Behavioral evidence may show whether response time is associated with:

  • Conversion

  • Appointment completion

  • Abandonment

  • Repeat purchase

  • Complaint rate

  • Customer retention

Customer behavior is especially important because internal process completion does not automatically mean that the process created the intended customer result.

A company may report that 98% of inquiries received an automated acknowledgment.

Customers may still abandon the process because they did not receive a useful answer, clear next step, appointment option, or personalized response.

Customer behavior can reveal whether the process is producing the outcome leadership cares about.

It does not automatically explain why the behavior occurred.

A customer may abandon because of:

  • Slow follow-up

  • Price

  • Service availability

  • Location

  • Competitor response

  • Confusing information

  • Personal circumstances

  • Poor fit

  • Timing

  • A change in need

Behavioral patterns should therefore be analyzed alongside process and customer information.

Customer complaints also require caution.

Complaints may identify serious failures and should not be dismissed.

But complaints usually lack a denominator.

Ten complaints could represent:

  • Ten failures among 100 customers

  • Ten failures among 100,000 customers

  • A concentrated problem in one location

  • A widespread problem reported by only a small proportion of affected customers

Complaints illustrate experience.

They do not automatically establish prevalence.

6. Employee Interviews

Employees often possess critical information that formal systems do not capture.

They know:

  • Which steps are routinely skipped

  • Which information is usually missing

  • Which approvals create delays

  • Which reports are inaccurate

  • Which system fields are misunderstood

  • Which workarounds keep the process functioning

  • Which customer situations create exceptions

  • Where responsibilities overlap

  • Where ownership is unclear

  • Which changes have already been attempted

  • Which problems are visible only during certain shifts or periods

Employee interviews can explain how and why the process behaves as it does.

They are particularly valuable for discovering what evidence should be collected next.

An employee might state:

We do not rely on the CRM notification because it arrives before the inquiry has been qualified. We wait for the intake coordinator’s email instead.

That statement does not prove how frequently the behavior occurs.

But it reveals a possible disconnect between the formal system and the actual coordination mechanism.

The diagnostic can then test it using:

  • Notification logs

  • Assignment times

  • Email records

  • Direct observation

  • A sample of inquiry histories

  • Interviews with other employees

Interviews are evidence, but they are self-reported evidence.

They may be affected by:

  • Memory

  • Fear of blame

  • Desire to protect coworkers

  • Desire to protect management

  • Personal incentives

  • Frustration

  • Limited visibility outside the employee’s role

  • Recent memorable events

  • Differences in terminology

  • Different interpretations of the same policy

  • Beliefs about what the interviewer wants to hear

Interview conditions matter.

Employees are more likely to provide useful information when:

  • The purpose of the diagnostic is explained

  • The discussion is not treated as a disciplinary interrogation

  • Questions are neutral

  • Employees can speak without direct supervision present

  • Confidentiality boundaries are explained accurately

  • The interviewer asks for examples and process details

  • Statements are checked against other sources

Compare these questions:

Why are employees failing to use the CRM correctly?

and:

Walk me through what happens from the moment an inquiry arrives until the first customer response is completed.

The first question assumes failure and assigns responsibility.

The second asks the employee to describe the process.

Interview evidence is strongest when the person has direct knowledge and the findings are corroborated through records, observation, or other interviews.

7. Management Interpretation

Management interpretation includes leadership’s explanation of:

  • What the process is intended to accomplish

  • What management believes is happening

  • Which standards apply

  • Which outcomes matter

  • Why prior decisions were made

  • Which constraints exist

  • Which risks are acceptable

  • Which changes are feasible

  • Which business consequences are material

Management interpretation is necessary.

Leaders often possess information employees do not have, such as:

  • Financial constraints

  • Contractual obligations

  • Strategic priorities

  • Customer commitments

  • Regulatory exposure

  • Previous failed initiatives

  • Vendor limitations

  • Planned organizational changes

  • Capacity decisions

  • Reasons for current policies

But management interpretation should not be treated as the highest form of operational evidence merely because it comes from authority.

Managers may be distant from the daily process.

Their understanding may be filtered through:

  • Summary reports

  • Departmental incentives

  • Prior decisions

  • Incomplete escalation

  • Confirmation bias

  • Sunk costs

  • Organizational politics

  • A desire to demonstrate control

  • A desire to assign responsibility

  • Different definitions of success

Management may describe the intended process accurately while misunderstanding the actual process.

For example:

Every inquiry is automatically routed to the next available salesperson.

System and observation evidence may show that inquiries are technically assigned automatically but remain invisible until an intake employee manually changes the record status.

Management’s statement is still useful.

It identifies the expected process and reveals the gap between design and operation.

Management interpretation should therefore be treated as:

  • Context

  • A source of hypotheses

  • A statement of intent

  • A source of decision criteria

  • One perspective to be tested against the operating evidence

It should not be dismissed.

It should also not be accepted without examination.

8. Anecdotes and Assumptions

Anecdotes are individual stories or isolated events.

Assumptions are beliefs accepted without adequate verification.

Examples include:

  • A customer complained that no one followed up.

  • One employee said the software is slow.

  • The owner believes younger customers prefer text messages.

  • A manager remembers that performance was better before the CRM changed.

  • Everyone knows the scheduling department is understaffed.

  • The new employee seems less productive.

  • Customers probably leave because competitors are cheaper.

  • AI should be able to automate the process.

Anecdotes and assumptions are not worthless.

They often identify issues that deserve investigation.

A single serious incident may reveal:

  • A safety risk

  • A compliance failure

  • A customer-service breakdown

  • A missing control

  • A process exception

  • A previously unknown vulnerability

But an anecdote does not establish:

  • Frequency

  • Prevalence

  • Typicality

  • Causation

  • Financial impact

  • Whether the event represents the broader population

A useful principle is:

Anecdotes generate questions. They do not settle them.

The statement:

A customer waited three days for a callback.

is an observation about one case.

It can lead to questions:

  • How often does this happen?

  • Which channel did the inquiry enter?

  • When was it assigned?

  • Was the record complete?

  • Did anyone attempt contact?

  • Was service capacity available?

  • Is the event part of a larger pattern?

The anecdote becomes the starting point for evidence collection.

It should not become the entire case for a major operational change.

The Hierarchy Does Not Replace Judgment

The evidence hierarchy helps organize the investigation.

It does not eliminate the need for judgment.

Consider two sources:

Source A

A CRM export containing 50,000 records.

However:

  • 35% of records have missing status information

  • Employees share user accounts

  • Historical timestamps were overwritten during a migration

  • Follow-up outside the CRM is not recorded

  • Definitions changed halfway through the period

Source B

A structured review of 150 randomly selected inquiries using original emails, telephone records, CRM activity, and employee confirmation.

Source A is higher in the general hierarchy because it is direct system data.

Source B may be more reliable for the specific question.

The diagnostic should evaluate evidence based on its actual quality, not merely its category.

Eight Dimensions of Evidence Quality

Regardless of source, operational evidence should be evaluated across several dimensions.

1. Relevance

Does the evidence directly address the problem or hypothesis?

A report may be accurate but irrelevant.

Total sales volume does not directly test whether delayed assignment causes slow follow-up.

2. Accuracy

Does the information correctly represent what occurred?

Accuracy may be affected by:

  • Entry errors

  • Incorrect formulas

  • Misclassification

  • Shared accounts

  • Faulty integrations

  • Improper exclusions

  • Incorrect timestamps

3. Completeness

Does the source contain the necessary records and fields?

Missing data may be concentrated in the transactions where the process failed.

That can bias the result.

4. Applicability

Is the evidence appropriate for the decision being made?

Data from one location, service, or period may not apply to the entire organization.

5. Representativeness

Does the sample reflect the relevant population and normal operating conditions?

Reviewing only complaints, escalations, or successful transactions creates a distorted picture.

6. Directness

How close is the source to the event or condition being evaluated?

Direct logs, records, inspection, and observation generally require less interpretation than secondhand reports.

7. Independence

Does the evidence come from a separate source, or are several reports built from the same underlying data?

Three dashboards generated from the same database are not three independent sources.

They are three presentations of one source.

8. Traceability and Reproducibility

Can another qualified person understand:

  • Where the data came from

  • How it was defined

  • What was included

  • What was excluded

  • How the calculation was performed

  • Whether the result can be reproduced

A conclusion that depends on undocumented spreadsheet changes or personal memory is difficult to verify.

EVIDENCE-QUALITY PRINCIPLE

A source is not strong merely because it is quantitative.

A number can be inaccurate, incomplete, irrelevant, unrepresentative, or generated from an unstable definition.

Evidence quality depends on how the information was created, controlled, selected, interpreted, and connected to the question.

Triangulation: Why One Source Rarely Proves Root Cause

Triangulation means examining the same operational problem through multiple sources or methods.

The purpose is not simply to collect more information.

The purpose is to determine whether different sources:

  • Converge on the same explanation

  • Add complementary information

  • Reveal different conditions across groups or periods

  • Contradict one another

  • Expose weaknesses in the measurement system

  • Identify a more complex causal structure

A process leaves evidence in several places.

It creates:

  • Digital footprints in system logs

  • A documentary trail in transaction records

  • Observable behavior in the workplace

  • Time patterns in queues and handoffs

  • Customer outcomes

  • Employee explanations

  • Management interpretations

Each source shows a different part of the journey.

No single footprint reconstructs the entire route.

Convergence

Convergence occurs when independent sources show a compatible pattern.

For example:

  • System logs show that late inquiries are assigned several hours after receipt.

  • Observation shows the intake coordinator batches assignments twice per day.

  • Time studies show that most elapsed time occurs before assignment.

  • Employee interviews confirm that only one role can complete qualification and routing.

Together, these sources provide stronger support for an assignment bottleneck than any source alone.

Complementarity

Sources may contribute different types of evidence without duplicating one another.

For example:

  • System data shows how often delays occur.

  • Observation shows where delays arise.

  • Interviews explain why the workaround exists.

  • Customer behavior shows whether the delays matter to conversion.

  • Financial analysis estimates the potential business exposure.

Each source answers a different question.

Together, they create a more complete operational explanation.

Contradiction

Contradiction is not necessarily a failure.

It is often diagnostically valuable.

Suppose:

  • The dashboard reports 96% on-time completion.

  • Transaction records show that many tasks were closed before customer contact.

  • Employees explain that tasks are closed to prevent overdue alerts.

  • Customer records show that actual follow-up occurs later.

The sources disagree because “task closed” is not equivalent to “customer contacted.”

The contradiction reveals a measurement-definition problem and possibly an incentive or workflow problem.

The correct response is not to choose the source management prefers.

The correct response is to determine why the sources disagree.

TRIANGULATION PRINCIPLE

Agreement across weak sources does not automatically create strong evidence.

Disagreement across credible sources should not be averaged away.

The diagnostic must determine whether the difference reflects data quality, population differences, process variation, changed definitions, or an incomplete explanation.

Worked Example: What the Evidence Actually Shows

Article 2 developed the following problem statement:

During the 90 days ending June 30, 2026, 290 of 420 qualified inquiries—69%—received a documented personalized follow-up within one business day, compared with the company’s target of 95%, creating a gap of 26 percentage points and an operational shortfall of 109 timely follow-ups.

Management’s initial explanation is:

Sales employees are ignoring CRM notifications.

Article 1’s framework requires competing hypotheses.

Possible explanations include:

Hypothesis A: Employees are ignoring assigned inquiries.

Expected pattern:

  • Inquiries are assigned promptly.

  • Employees can see the assignment.

  • Late cases remain untouched after assignment.

  • Delay varies materially by assigned employee.

Hypothesis B: Inquiries are assigned too late.

Expected pattern:

  • Late follow-up is preceded by late qualification or assignment.

  • Most elapsed time occurs before the inquiry reaches the responsible employee.

  • Delay is concentrated during certain intake periods.

Hypothesis C: Follow-up occurs but is not documented.

Expected pattern:

  • External email, telephone, or text records show contact not recorded in the CRM.

  • Employees describe using communication channels outside the approved record.

  • Customer responses occur despite missing CRM activity.

Hypothesis D: Workload exceeds available capacity during peak periods.

Expected pattern:

  • Delays increase when inquiry volume exceeds available staffing.

  • Queue time increases more than task-processing time.

  • Performance improves during lower-volume periods or when additional capacity is available.

Hypothesis E: Employees delay contact when service capacity is unavailable.

Expected pattern:

  • Delayed inquiries are concentrated in services, dates, or locations with limited availability.

  • Employees report waiting for scheduling information before contacting customers.

  • Fast assignment does not consistently produce fast follow-up when appointment capacity is constrained.

The diagnostic then examines several evidence sources.

Direct System Data

The system history shows:

  • 76% of late-follow-up cases were assigned more than six business hours after receipt.

  • Once assigned, most inquiries were opened by the employee within one hour.

  • Delay varied more by time of receipt than by assigned salesperson.

  • Inquiries received after 2:00 p.m. were substantially more likely to miss the standard.

This evidence weakens the theory that sales employees are generally ignoring promptly assigned records.

It strengthens the possibility of an intake or assignment bottleneck.

Transaction Records

A sample of late cases is compared across CRM records, email, and telephone logs.

The review finds:

  • Some inquiries received contact outside the CRM.

  • Most late inquiries had no evidence of contact before assignment.

  • Several records lacked the information needed to determine the requested service.

  • Some inquiries were reassigned several times.

This shows that documentation problems exist, but they do not explain the entire performance gap.

Direct Observation

Observation of the intake workflow shows:

  • Website inquiries enter a shared inbox.

  • The CRM receives the record immediately.

  • The intake coordinator reviews each inquiry manually.

  • Qualification requires checking service area, requested service, and appointment availability.

  • The coordinator batches assignments near midday and late afternoon.

  • Only the coordinator and one manager can finalize the assignment status.

The “automatic assignment” described by management is not fully automatic.

The system creates the record automatically.

The inquiry does not become operationally available to the salesperson until manual qualification and assignment are completed.

Time Study

The time study finds:

  • Average qualification touch time: four minutes

  • Average assignment touch time: one minute

  • Average wait before qualification: three hours

  • Average wait after qualification but before assignment: two hours

  • Average employee response after assignment: 48 minutes

Most of the elapsed time occurs before the responsible employee receives the inquiry.

This evidence further weakens the initial theory.

Customer Behavior

The customer analysis shows:

  • Inquiries receiving follow-up within one business day convert more frequently.

  • The difference is larger for services with near-term appointment availability.

  • When appointment availability is limited, conversion remains low even with timely follow-up.

This suggests that response time matters but may interact with service capacity.

It does not support a simple one-cause explanation.

Employee Interviews

Employees report:

  • They do not act on the initial CRM notification because the inquiry may not yet be qualified or assigned.

  • They wait for the assignment email.

  • Some contact activity occurs through personal or mobile telephone systems and is entered later.

  • They are sometimes reluctant to contact customers when they cannot offer an appointment.

These statements explain several patterns observed elsewhere.

They also identify possible documentation and service-capacity issues.

Management Interpretation

Management believed that:

  • The CRM automatically assigns every inquiry.

  • Sales employees receive immediate notification.

  • The primary problem is insufficient employee accountability.

The evidence shows that the initial system event occurs automatically, but the usable assignment depends on manual qualification and routing.

Management’s interpretation was understandable but incomplete.

Triangulated Interpretation

The combined evidence suggests:

  1. The largest delay occurs before sales employees receive a usable assignment.

  2. Manual qualification and batched routing create a queue.

  3. Access permissions create a single-point dependency.

  4. Some follow-up occurs outside the approved documentation process.

  5. Service availability influences both employee behavior and customer conversion.

  6. Employee accountability may still matter in individual cases, but it does not appear to be the primary explanation for the overall gap.

This is stronger than concluding:

Employees are ignoring the CRM.

It is also more precise than concluding:

The CRM is the problem.

The system configuration, qualification workflow, access structure, documentation practices, staffing pattern, and service-capacity information may all contribute.

The next step should not automatically be full CRM replacement.

A more defensible next step might be a controlled pilot that:

  • Routes qualified inquiries continuously rather than in batches

  • Expands appropriate assignment permissions

  • Clarifies which system event requires employee action

  • Integrates or standardizes contact documentation

  • Provides employees with current service-capacity information

  • Measures assignment time, follow-up time, documentation completeness, conversion, rework, and workload

If performance improves during the pilot without creating unacceptable new problems, the organization will have stronger evidence that the countermeasure addresses the supported causes.

Triangulation Does Not Automatically Establish Causation

Multiple sources can create a credible causal explanation.

They may still stop short of proving that changing one factor will produce the desired result.

System data may show that delayed assignment and delayed follow-up occur together.

Observation may show how batching creates the delay.

Interviews may confirm why the batching occurs.

Customer behavior may show lower conversion in delayed cases.

That creates a strong operational case.

But the most direct test may still be to change the assignment process for a defined group and observe whether:

  • Assignment time improves

  • Follow-up time improves

  • Customer conversion changes

  • Employee workload remains manageable

  • Rework does not increase

  • Service quality does not decline

The diagnostic identifies the most credible explanation and the most appropriate countermeasure to test.

Implementation evidence confirms whether the change produces the expected result.

State What Each Source Can and Cannot Establish

A disciplined evidence plan should specify the limits of each source.

Direct system data can often establish:

  • Frequency

  • Timing

  • Sequence

  • Distribution

  • Status

  • Volume

  • Patterns across transactions

It may not establish:

  • Whether the recorded action was performed correctly

  • Why the action occurred

  • Whether work happened outside the system

  • Causation

Transaction records can often establish:

  • What was documented

  • Which information was available

  • Whether required records exist

  • How an individual case progressed

They may not establish:

  • Whether the documentation is complete

  • Whether the process was performed as recorded

  • How representative the case is

Observation can often establish:

  • How work is performed

  • Which workarounds exist

  • Where waiting or rework occurs

  • How employees interact with systems

It may not establish:

  • Long-term frequency

  • Typical performance across all people or periods

  • Customer or financial impact

Time studies can often establish:

  • Duration

  • Wait time

  • Touch time

  • Variation

  • Queue behavior

They may not establish:

  • The cause of every delay

  • Whether the observed period is representative

  • Whether faster work produces a better outcome

Customer behavior can often establish:

  • What customers did

  • Where they abandoned

  • Whether outcomes differ across conditions

It may not establish:

  • Why customers behaved that way

  • Which internal factor caused the outcome

Employee interviews can often establish:

  • Perceived barriers

  • Workarounds

  • Process history

  • Definitions

  • Possible explanations

  • Information not captured elsewhere

They may not establish:

  • Prevalence

  • Objective frequency

  • Causation

  • Whether the account is representative

Management interpretation can often establish:

  • Intended standards

  • Strategic priorities

  • Decision criteria

  • Organizational constraints

  • Historical context

It may not establish:

  • The actual daily workflow

  • The frequency of specific failures

  • The primary root cause

Anecdotes can often establish:

  • That a specific event occurred

  • That a risk or failure may deserve investigation

They may not establish:

  • How often it occurs

  • How material it is

  • Whether it represents the broader process

  • Why it occurred

Build the Evidence Plan Before Collecting the Evidence

An evidence plan should be developed before the organization decides which solution it wants.

For each hypothesis, document:

The claim

What explanation is being tested?

The expected pattern

What should be visible if the explanation is substantially correct?

The primary evidence

Which source most directly tests the claim?

Corroborating evidence

Which independent or complementary sources can support or challenge the finding?

Reliability risks

What could make the evidence inaccurate, incomplete, biased, or unrepresentative?

Disconfirming evidence

What result would weaken or overturn the explanation?

Decision use

What action would the evidence justify?

For example:

Hypothesis: Delayed manual assignment is the primary contributor to late customer follow-up.

Expected pattern:

Late-follow-up cases should have materially longer assignment delays than timely cases. Once assigned, employee response should generally occur within the required remaining period.

Primary evidence:

CRM receipt, qualification, assignment, and contact timestamps.

Corroborating evidence:

Direct observation, time study, transaction review, and employee interviews.

Reliability risks:

Missing timestamps, overwritten status histories, contact outside the CRM, and inconsistent qualification definitions.

Disconfirming evidence:

Assignment occurs promptly in most late cases, or late follow-up continues after assignment delay is removed.

Decision use:

Determine whether to pilot continuous routing, revise permissions, or investigate employee-level response behavior instead.

This structure reduces the risk of collecting only the information that supports the preferred solution.

More Evidence Is Not Always Better

An operations diagnostic should collect enough evidence to support the decision.

It should not collect every available record merely because the information exists.

Unnecessary collection creates:

  • Cost

  • Delay

  • Privacy exposure

  • Security risk

  • Analytical noise

  • Scope expansion

  • Greater opportunity for misinterpretation

The appropriate amount of evidence depends partly on the decision.

A low-cost, reversible pilot may require less certainty than:

  • Replacing a major software platform

  • Eliminating positions

  • Restructuring a department

  • Changing customer pricing

  • Making a regulatory representation

  • Altering a safety-critical process

  • Committing substantial capital

  • Changing a contractual service standard

The higher the cost, risk, and irreversibility of the decision, the stronger the evidentiary requirement should be.

Translate Evidence Strength Into the Appropriate Action

Evidence does not produce only two possible outcomes: act or do nothing.

Different evidence conditions justify different responses.

Weak or unreliable evidence

Appropriate response:

  • Repair the measurement system

  • Clarify definitions

  • Collect missing records

  • Improve documentation

  • Narrow the question

  • Avoid major decisions

Conflicting evidence

Appropriate response:

  • Investigate why sources disagree

  • Segment by process, employee, channel, location, or period

  • Test whether definitions changed

  • Examine data-quality problems

  • Revise the hypothesis

Suggestive but incomplete evidence

Appropriate response:

  • Run a limited pilot

  • Collect additional targeted evidence

  • Monitor outcome, process, and balancing measures

  • Avoid full-scale implementation

Strong convergent evidence

Appropriate response:

  • Implement a defined countermeasure

  • Establish ownership

  • Measure implementation fidelity

  • Monitor the original performance gap

  • Watch for unintended consequences

Evidence that the countermeasure changed the result

Appropriate response:

  • Standardize the improved process

  • Train affected employees

  • Update documentation and controls

  • Scale deliberately

  • Continue monitoring for regression

The purpose of evidence is not to create permanent analysis.

It is to match the strength of the management action to the strength of the information.

The Eight-Question Management Test

Before leadership accepts a claimed root cause or approves a major operational change, it should be able to answer these questions in writing:

  1. What defined performance gap are we trying to explain?

  2. What competing hypotheses could reasonably produce the same observed condition?

  3. What is the strongest direct evidence available for or against each hypothesis?

  4. Have the system data and transaction records been assessed for accuracy, completeness, applicability, definitions, and control weaknesses?

  5. What did direct observation, time analysis, customer behavior, and employee interviews add that the system data could not show?

  6. Are the findings representative of the relevant transactions, employees, locations, channels, and time periods?

  7. Where do the sources converge, complement one another, or contradict one another—and have the contradictions been explained?

  8. What action is justified by the current strength of evidence: repair measurement, investigate further, run a pilot, implement a scoped change, or scale a proven improvement?

If leadership cannot answer these questions, the organization may possess information without possessing a defensible diagnosis.

Conclusion: Root Cause Is Earned Through Converging Evidence

Operational problems produce many explanations.

The dashboard says employees are late.

Employees say the system is slow.

Management says accountability is weak.

Customers say communication is unclear.

The software vendor says configuration is the problem.

Each explanation may contain part of the truth.

The purpose of an operations diagnostic is not to choose the most confident speaker.

It is to determine which explanation best fits the evidence.

Article 2 defined the measurable gap.

Article 1 established the need for competing, testable hypotheses.

Article 3 establishes how those hypotheses should be evaluated.

The full diagnostic chain is:

Defined gap → competing hypotheses → evidence hierarchy → triangulation → supported causes → tested countermeasures → measured results → operational control

Direct system data may show the scale and timing of the problem.

Transaction records may reconstruct what happened.

Observation may reveal the actual workflow.

Time studies may show where work waits.

Customer behavior may show whether the process affects outcomes.

Employee interviews may explain hidden constraints and workarounds.

Management interpretation may establish intent, priorities, and decision criteria.

Anecdotes may identify where investigation should begin.

Each source contributes a different piece.

A credible conclusion emerges when the pieces form a coherent explanation, withstand contradictory evidence, and predict what should happen when the process changes.

OPERATING RULE

No major process, staffing, software, AI, training, pricing, or policy change should be approved solely on the basis of one dashboard, one interview, one complaint, one management belief, or one isolated correlation.

Leadership should require evidence from multiple relevant sources, assess the reliability and limitations of each source, investigate contradictions, and match the scale of the decision to the strength of the evidence.

When Structured Diagnosis Becomes Appropriate

When an established organization has a recurring operational problem with a meaningful customer, financial, employee, capacity, service-delivery, or management impact—and the available information is incomplete, conflicting, or insufficient to establish the cause—Angulo & Morsa’s Operations Diagnostic examines the actual workflow, available system data, transaction records, direct observations, employee and management input, customer outcomes, measurement limitations, and competing explanations before recommendations are made.

[Learn about the Angulo & Morsa Operations Diagnostic]

Related Reading

[Read Article 1: Stop Guessing at Root Causes]

[Read Article 2: How to Write an Operational Problem Statement]

[Return to The Legacy Perspective Knowledge Library]

Educational and Professional Disclaimer

This article is provided for general educational and informational purposes.

It does not constitute legal, tax, accounting, financial, investment, human-resources, cybersecurity, regulatory, engineering, statistical, audit, investigation, or other professional advice.

Examples and calculations are illustrative and may not apply to a specific organization.

Operational evidence depends on data quality, definitions, access, sampling, internal controls, process conditions, employee and customer privacy, and organizational context. Certain matters may require legal counsel, licensed professionals, regulatory specialists, auditors, cybersecurity professionals, statisticians, or other qualified experts.

Use of this article does not create a consultant-client or other professional relationship with Angulo & Morsa Legacy Consulting LLC.

References

American Society for Quality. (n.d.). DMAIC process: Define, measure, analyze, improve, control.

Besculides, M., Zaveri, H., Farris, R., & Will, J. (2006). Identifying best practices for WISEWOMAN programs using a mixed-methods evaluation. Preventing Chronic Disease, 3(1).

Government Accountability Office. (2019). Assessing data reliability (GAO-20-283G).

Government Accountability Office. (2024). Government Auditing Standards: 2024 revision (GAO-24-106786).

Institute for Healthcare Improvement. (n.d.). Model for Improvement: Establishing measures.

Lean Enterprise Institute. (n.d.). Genchi genbutsu.

National Institute of Standards and Technology. (2025). Measurement assurance.

© 2026 Angulo & Morsa Legacy Consulting LLC. All rights reserved.

Previous
Previous

Article 4: Map the System Before Changing the Process

Next
Next

Article 2: Stop Guessing at Root Causes