Fault diagnosis and detection
Repairing is not swapping parts until it works. It is reducing uncertainty with measurements that rule out half of the problem at a time. A technician with a method and a multimeter finds more faults than one without a method and the best instruments.
01The method
- Listen to the whole symptom and write it down: what it does, what it doesn't do, since when, what changed beforehand. Half of all faults get narrowed down in that conversation.
- Reproduce the fault. What can't be reproduced can't be verified as repaired.
- Visual and smell inspection: burnt, swollen, loose, corroded, overheated. It is free and solves a great deal.
- Hypothesis: what could produce exactly that symptom.
- A measurement that rules out, not one that confirms: look for the point where the hypothesis and its opposite give different results.
- Repair, verify and record, including the root cause: if a transistor burned out, why it burned out.
Substitution only proves something when the replacement is identical and good. Swapping without measuring has three costs: it uses up spare parts, it can burn the new part for the same reason the old one burned, and —worst of all— it hides the root cause. A unit that comes back to the shop every two months is always that: the effect was replaced, not the cause.
02Bisection search
It is the most powerful technique and the simplest. If the signal goes in fine at one end and comes out wrong at the other, you measure in the middle: that rules out half of the circuit in one go.
| Technique | What it involves | When it fits |
|---|---|---|
| Bisection | Measure in the middle of the chain and rule out half | Whenever there is a chain of stages: audio, communication, control |
| Signal injection | Feed a known signal in at an intermediate point and see whether it appears at the output | When the original signal doesn't exist or can't be reproduced |
| Signal tracing | Follow the chain with the oscilloscope from the input | When the signal is present and degrades somewhere along the way |
| Comparison | Measure the same points on a working unit | Identical units across the plant. It is very fast and very reliable |
| Cold and heat | Freeze spray or hot air on a suspect component | Faults that appear or disappear with temperature |
| Controlled tapping | Press on components and connectors with the equipment running | Cold solder joints and loose contacts |
03What each instrument reveals
| Instrument | What it finds | What it does NOT see |
|---|---|---|
| Multimeter | Continuity, quiescent voltages, resistances, open or shorted semiconductors | Anything that happens fast: ripple, oscillations, pulses |
| Oscilloscope | Waveform, ripple, noise, timing, parasitic oscillations | Faults that occur once an hour, unless it is left capturing |
| ESR meter | Worn-out electrolytic capacitors, in circuit | Internal short circuits of other components |
| Clamp meter | Current draw without opening the circuit, imbalances between phases | Very small currents |
| Thermal imaging camera | The hot spot before it fails: loose terminals, overloaded components | Cold faults (an open circuit doesn't heat up) |
| Megohmmeter | Degraded insulation in motors and cables | It is not used on electronics: the test voltage destroys it |
Before opening anything, measure how much current the unit draws and compare it with normal. A very high draw points to a short circuit; a very low one, to something not starting; a normal one with the unit not working, to a signal or control problem. Three very different possibilities told apart by a single measurement.
04Typical faults by family
- Electrolytic capacitors with high ESR: it is the most common fault of all.
- Shorted rectifier diodes, which often take the fuse with them.
- Open startup resistor in switching supplies.
- Degraded feedback optocoupler: the output voltage drifts upward.
- Short-circuited semiconductor, almost always with another cause behind it: overload, insufficient heat dissipation, missing snubber.
- Open gate resistors.
- Burnt traces and cold solder joints from thermal cycling.
- Missing decoupling: random lock-ups.
- Floating inputs reading noise.
- Badly built reset, crystal that doesn't start.
- Program that hangs, with no watchdog to recover it.
- Missing termination or common ground.
- Cable running next to a power cable.
- Baud rate set wrong.
- A node left transmitting that blocks the bus.
05Intermittent faults
These are the ones that consume the most time, because by the time the technician arrives the equipment works. The strategy is different: instead of looking for the fault, you have to set a trap for it.
- Log: leave a data logger or a microcontroller recording voltages, current and temperature with date and time. When the fault occurs, the exact moment will be on record, along with what was happening around it.
- Correlate: does it always happen at the same time of day? When another motor starts? When it rains? When it warms up? That correlation is often worth more than any measurement.
- Provoke it: heat, cold, vibration, moving the wiring with the equipment running.
- The multimeter's MIN/MAX and the digital oscilloscope's event capture: they keep a record of the peak that occurred while nobody was watching.
A machine stops with no apparent pattern. Instead of replacing the drive:
- The date and time of each stop are noted down for a month.
- The correlation shows up: always between 1 and 2 p.m., the time of highest consumption in the plant.
- The line voltage is logged: it drops to 195 V at that time.
- The drive shuts off on undervoltage, exactly as it is designed to do.
The fault was not in the equipment. No measurement taken at 9 in the morning would have found it.
06Measuring with the equipment energized
- Never alone. Always with another person present who knows how to cut the power.
- Instrument and leads of the right category (digital instruments), in good condition and checked before and after.
- Isolation transformer for the oscilloscope, or a differential probe. Never connect the oscilloscope ground to a live point.
- Only one hand working; keep the other out. No rings, watch or metal bracelets.
- Insulating footwear, dry floor, personal protective equipment.
- Filter capacitors discharged and verified before touching: in a switching supply they stay above 300 V for minutes.
07In the lab
The instructor introduces a fault into a multi-stage amplifier chain. One group looks for it stage by stage from the input and another by bisection, and both are timed. It is repeated with the fault in another stage. The time difference is the point of the lab.
On identical boards, seed five different faults: electrolytic capacitor with high ESR, open resistor, cold solder joint, shorted transistor and loose connector. Each group diagnoses one and writes the report: symptom, measurements, conclusion and root cause. Then they rotate.
With two identical units, one good and one faulty, measure the same points on both and note the differences. It is the fastest method when it is possible, and it teaches where to look in the future.
Seed a cold solder joint and look for it with the equipment running: gentle tapping, pressure with a wooden stick, freeze spray and hot air. Note which of the four methods revealed it and why.
08Common mistakes
| Technician's mistake | Consequence |
|---|---|
| Replacing the burnt component and returning the unit | It comes back to the shop: nobody looked for why it burned. |
| Not asking what happened before the fault | The most valuable clue is lost: “water got into it,” “there was a thunderstorm,” “it was moved.” |
| Measuring to confirm the hypothesis | You find what you look for. You have to measure to rule out. |
| Taking everything apart before measuring | The state in which it failed is lost, and sometimes the fault disappears by itself. |
| Trusting your eyes to judge an electrolytic capacitor | Many are worn out without looking swollen. An ESR meter is needed. |
| Measuring semiconductors in circuit and drawing conclusions | The rest of the circuit distorts the measurement: you have to lift a lead or compare with a good one. |
| Working with the equipment energized when it isn't necessary | An avoidable risk. Most measurements are made with the power off. |
| Not recording anything | Next time you start from zero, and repeat faults go undetected. |
09Self-assessment
How many measurements does bisection need to isolate the faulty stage among eight?
Three: each measurement rules out half (8 → 4 → 2 → 1). Going stage by stage could take eight.
Why should a measurement aim to rule out rather than confirm?
Because a measurement that only confirms what was already believed adds no information. A useful one separates two hypotheses: it gives one result if it is one and another if it is the other.
The equipment draws far more current than normal. What is suspected?
A short circuit or a shorted semiconductor. A very low draw would point to something not starting, and a normal draw with the equipment idle, to a signal or control problem.
Why isn't looking at an electrolytic capacitor enough to judge it?
Because it can be worn out —with very high ESR— without looking swollen or leaking. The only way to know is to measure the ESR, and it can be done with the component on the board.
What precaution is required when measuring with the oscilloscope on a circuit connected to the mains?
Use an isolation transformer or a differential probe. The oscilloscope ground is tied to earth: connecting it to a live point causes a dead short circuit.
What is the best strategy for an intermittent fault?
Log and correlate: keep a time-stamped record of variables and look for what the fault coincides with. And provoke it with temperature, vibration or movement of the wiring.
The power transistor was replaced and failed again a week later. What was missed?
Looking for the root cause: insufficient heat dissipation, missing snubber, overload, slow driver or a fault in the load. The transistor was the effect, not the cause.
What does a thermal imaging camera reveal that a multimeter does not?
The hot spots before they fail: a loose terminal, an overloaded component, an unbalanced phase. It is the central tool of predictive maintenance.
Why is it worth measuring the same point on a working unit?
Because it gives the real correct value for that circuit, with its load and its conditions, without relying on assumptions or on documentation that may not match what is installed.
What should be written down at the end of a repair?
Symptom, measurements and values, component replaced, root cause, how the repair was verified, date and person responsible. Without that there is no history and no way to detect repeat faults.