Case 09 · Acoustic Diagnostics · Application Playbook

Acoustic AI by Site
Seven Deployment Cases for an 8-Mic RK3308 Array

All 7 of these sites are different problems — and not because the microphone changes. They are different because how long you must listen changes. Each case below is written the same way: the site, the signature, the interference, the mounting, the array, the decision, the output, and honestly where the numbers come from.

How to read this. These are engineering scenarios, not delivered projects. No client is named, and none of these sites has been commissioned by us. Each case is a design inference built from four graded kinds of input, labelled throughout: [chip] Rockchip RK3308 product page and datasheet; [physics] derived here from the speed of sound and array geometry; [industry] published practice in vibration monitoring, acoustic emission and structural health monitoring; [engineering] our own design conclusions drawn from the three above. Where a value can only be meaningful after it is measured on site, we say so rather than invent one.

Before the cases

Three Things Decide Everything

Frequency band, resolution and pipeline structure are set by three separate things. Confusing them is the most common way an acoustic monitoring project fails at the specification stage.

Aperture sets the band

An 8-element array holds its directivity over only 1.8 octaves — and that span does not depend on the element spacing. Spacing slides the window; it never widens it. Widening needs more elements: 2.9 octaves at 16, 4.0 at 32.

Window length sets the resolution

Frequency resolution is Δf = 1/T, where T is the length of one analysis window. To resolve 0.1 Hz you must watch for at least 10 s; to resolve 20 kHz, 50 µs. That is the range a single 8-mic front end is being asked to span.

Time scale sets the pipeline

Across the 22 scenarios the judgement time scale runs from 1 ms to 3600 s — about 6.6 orders of magnitude (3600000×). It is this spread, not the band, that forces separate analysis paths.

1 ms1 s1 min1 h5674Judgement time scale: 1 ms to 3600 s = 3,600,000 × = 6.6 orders of magnitudeBar width is fixed: the gaps between classes are empty, not gradual.
Resolution, derived

The Minimum Observation Window, per Frequency

Δf = 1/T is not a tuning choice, it is a floor: you cannot distinguish two features closer than 1/T apart. The right-hand column names the case in this article that sets each requirement.

Frequency to resolveMinimum windowWhich case sets this
0.1 Hz10 sTower modal drift — Case 2
1 Hz1 sBridge span response — Case 5
20 Hz50 msFall impact, debris-flow onset — Cases 5, 4
100 Hz10 msRoad and rail pass-by — Case 5
1 kHz1 msDry-boil, fan blade pass — Cases 6, 3
8 kHz125 µsGlass break — Case 6
20 kHz50 µsRockfall, gear mesh — Cases 4, 1
50 kHz20 µsPipe-leak ultrasound — Case 6

Each row is the shortest window in which that frequency is resolvable at all. In practice you take several windows' worth for a stable estimate — the numbers above are the floor, not the setting.

1 Hz100 Hz10 kHz100 s1 s10 ms100 µsfrequency to resolve (log)observation window (log)0.1 Hz → 10 s20 Hz → 50 ms1 kHz → 1 ms50 kHz → 20 µsresolvablebelow the resolution floor

The consequence for this whole article: a dense and a sparse aperture can sit on the same board and serve two different frequency bands at once. But no arrangement of apertures makes the window flexible — window length is set by the event, and the events here differ by 6.6 orders of magnitude. Bands can share an installation. Windows cannot.

The cases

Seven Sites, Seven Different Problems

Each case is laid out identically so they can be compared side by side. The eight fields are the same eight questions you would ask a systems engineer in a kickoff meeting.

Case 1 · Industrial site

Rotating & Reciprocating Machinery

A press shop or a pump hall: motors, gearboxes and a press sharing one floor. Mains power, an existing cabinet, a plant network — and a background that is loud, broadband and shift-dependent.

Site

Indoors, hard-surfaced, reverberant. Mains and network available, cabinet space exists, and maintenance staff already visit on a schedule. Background level is set by the machines themselves and changes with the shift.

What to listen for

Bearing defect tones and their sidebands against shaft harmonics, gear-mesh frequencies, and the impulsive signature of a press stroke. The signal of interest is rarely the loudest sound in the room — it is the part that changes.

Dominant interference

The machine's own tonal content. An incipient bearing fault contributes far less energy than the running tones, so broadband level monitoring saturates long before the defect becomes visible. Narrow-band tracking of a specific order is what works.

Mounting & protection

On the machine frame, or on a mast beside it, protected to IP65 and rated for the hall's ambient temperature. Not on the floor slab if the press is on it — the slab carries the press impulse and will dominate everything.

Array & sample rate

The group spans 10 Hz – 20 kHz, with a peak requirement of 48 kHz. No single aperture covers that: expect a dense aperture for the 1 kHz-and-up gear and press content, plus a wider one for the low-order shaft tones. This is the two-aperture pattern from the Tech Note, in its simplest form.

Edge decision

Baseline per machine and per operating state, then track order-band energy against that baseline. Bearing degradation is an hours-scale judgement; a press stroke is a milliseconds-scale event. Same sensor, two very different windows — see the two-window section below.

Output & response

Local closed loop. A graded alarm to the maintenance terminal plus a stored feature vector. No raw audio leaves the plant. Escalation to an engineer carries the trend, not the waveform.

Basis & uncertainty

Order and mesh frequencies come from geometry and speed, so they are computable from nameplate data. Bearing-defect band ratios are [industry] practice. Early-fault detectability and alarm thresholds are [engineering] and must be calibrated on the actual machine — we do not quote a universal threshold.

Scenarios in this group

ScenarioBand of interestSuggested sample rateJudgement time scaleDecision locus
Motor / pump — unbalance, misalignment, bearing10 Hz – 10 kHz≥ 24 kHzhoursOn device
Gearbox / reducer — mesh, tooth breakage100 Hz – 20 kHz≥ 48 kHzhoursOn device
Press / machine tool — impact, tool wear200 Hz – 20 kHz≥ 48 kHzmillisecondsOn device
Case 2 · Energy site

Wind Turbines, Towers & Transformers

Three different structures with one thing in common: the interesting frequencies are extremely low, the sites are remote, and nobody is standing there.

Site

An unattended outdoor site. A turbine tower, a transmission tower and a distribution transformer all present temperature extremes, continuous vibration and long maintenance intervals. Power is either available at the base or must be harvested.

What to listen for

Tower modal frequencies and their drift — a loss of stiffness shows up as a shift — plus blade-pass modulation, and transformer core and winding vibration at mains harmonics.

Dominant interference

Wind. It is the loudest thing on site and it is not the signal. Gust loading, blade swish and rain all modulate the same low-frequency band you are trying to watch, and none of them is stationary.

Mounting & protection

Inside the tower base or on the nacelle; on the transformer tank wall; on the tower legs. The enclosure must survive weather without a maintenance visit, so corrosion and cable entry matter as much as the sensor.

Array & sample rate

The band starts at 0.1 Hz – 10 kHz — below the reach of any compact aperture, so a large sparse aperture is mandatory. This is also the group that needs two sample rates in parallel: a normal rate for the acoustic content, plus a slow channel for the sub-hertz trend.

Edge decision

Modal tracking is a minutes-to-hours judgement on a long window. This group is local-trigger-plus-cloud-review: a single node cannot separate a stiffness change from a wind event on its own.

Output & response

Local trigger, plus a slow feature stream upward. At any useful sample rate the raw stream is enormous, so what leaves the site is modal frequencies, damping estimates and deviation trends — not audio.

Basis & uncertainty

Tower modal frequency depends on the real structure and its foundation, so it cannot be inferred from a datasheet and must be measured. Blade-pass and mains harmonics are computable from nameplate data. Damping estimation accuracy against a specific tower is [engineering].

Scenarios in this group

ScenarioBand of interestSuggested sample rateJudgement time scaleDecision locus
Wind turbine — main bearing, blade, yaw0.1 Hz – 10 kHz≥ 24 kHz + slow channelhoursTrigger + review
Tower / blade — modal and low-frequency response0.1 – 20 Hz≥ 100 HzminutesTrigger + review
Transformer / reactor — abnormal hum50 Hz – 5 kHz≥ 12 kHzhoursOn device
Case 3 · Livestock site

Livestock Houses: Fans, Animals & Feed Lines

A long, hard-surfaced shed: rows of ventilation fans at one end, animals in pens, an automatic feeding line — and a humidity, dust and wash-down regime that kills ordinary enclosures.

Site

A harsh indoor environment by any measure: humid, dusty, ammonia-bearing air, regular high-pressure wash-down, and a long reverberant volume. The uplink is often unreliable, which pushes decisioning to the edge by default.

What to listen for

Two very different things. First, fan and drive degradation — bearing, belt and imbalance. Second, animal vocalisation: distress calls, coughing, and the shift in a group's sound when something in the house changes.

Dominant interference

Reverberation. A hard shed with parallel walls is an acoustically hostile box; the direct sound from one pen is buried in the diffuse field. On top of that, the fans never stop, so there is no quiet reference to compare against.

Mounting & protection

Rigid mounting, positioned out of the fan airflow to avoid flow noise directly on the microphone, high enough to clear animals and wash-down spray. Sealing and cable routing matter more here than the sensor choice.

Array & sample rate

The group spans 20 Hz – 10 kHz with a peak requirement of 24 kHz. A mid aperture serves the fan content; animal vocalisation sits higher and needs less aperture but more channels, because separating pens is a spatial problem.

Edge decision

Fan degradation is a minutes-scale judgement against a per-fan baseline. Animal sound changes on a seconds scale. Both are local closed loop, because the shed may have no reliable uplink at all.

Output & response

A graded fan-maintenance alarm and an animal-welfare indicator, both computed locally. What leaves the site is the indicator and its confidence — not the audio.

Basis & uncertainty

Fan defect frequencies are [industry] practice and need a per-model baseline. Animal vocalisation classifiers need species-specific data, so we describe the method and do not claim an accuracy figure.

Scenarios in this group

ScenarioBand of interestSuggested sample rateJudgement time scaleDecision locus
Ventilation fan — bearing, belt, rotor lock20 Hz – 10 kHz≥ 24 kHzminutesOn device
Animal vocalisation — cough, alarm, farrowing50 Hz – 8 kHz≥ 24 kHzsecondsOn device
Feed line / scraper — jam and abnormal impact20 Hz – 8 kHz≥ 24 kHzminutesOn device
Case 4 · Geohazard site

Rivers & Valleys: Debris Flow and Rockfall

A mountain stream or a valley wall. No mains, no network, fully weather-exposed — and the event you care about may happen once a year, or once an hour.

Site

A remote outdoor site with no infrastructure. Power is solar or harvested, the uplink is intermittent at best, and the unit must be assumed unreachable for months. Snow, frost and flood all have to be designed for.

What to listen for

A change in the river's own sound — a debris flow raises the bed-load impact rate and lowers the spectral centroid — and the impulsive crack-and-roll of a rockfall starting on a slope.

Dominant interference

Rain and wind. Rain on the enclosure is broadband and impulsive, which makes it look a great deal like the events being hunted. This is the hardest false-positive problem in the whole set.

Mounting & protection

On a mast or a rock anchor, above flood level, aimed at the channel or the slope. The enclosure has to shed water and take an impact; the cabling has to survive frost and animal damage.

Array & sample rate

Flow-change judgement needs content across 10 Hz – 20 kHz, so wideband capture is warranted, and the rockfall is impulsive and needs the high end. Peak requirement 48 kHz.

Edge decision

The rockfall is a milliseconds-scale event — catch it with a short trigger window. The debris flow is a minutes-scale trend in continuous sound. Two windows at one site, which makes this group the clearest example of the two-stage structure.

Output & response

Local decision plus cloud explanation. The local node triggers immediately; the cloud is used to separate a real event from rain by looking at longer context and at neighbouring nodes. A false alarm here costs a response team, so the two-stage split is a safety feature rather than an optimisation.

Basis & uncertainty

The acoustic signature of bed-load transport is [industry] research rather than a settled standard. Thresholds for a specific channel require recording across seasons; we do not offer a universal threshold.

Scenarios in this group

ScenarioBand of interestSuggested sample rateJudgement time scaleDecision locus
River acoustics — discharge and turbulence shift20 Hz – 5 kHz≥ 12 kHzminutesOn device
Debris flow / flash flood — low rumble plus sustained scour10 Hz – 2 kHz≥ 6 kHzminutesDevice + cloud
Rockfall / slope collapse — impact transient and rolling rhythm100 Hz – 20 kHz≥ 48 kHzmillisecondsOn device
Case 5 · Infrastructure

Bridges, Roads, Rail, Cables & Buildings

Five infrastructure types that share one driver: traffic. A bridge deck, a road surface, a rail track, a stay cable and a tall building — and in every one of them, the excitation is also the interference.

Site

Outdoor, exposed, and in most cases impossible to close for installation. Access — not electronics — is the binding constraint on this group, and it shapes the whole deployment plan.

What to listen for

Vehicle pass-by signature (axle count and timing, hence speed), the structure's response to a known load, rail–wheel interaction, cable vibration and its damping, and in a building the difference between normal occupancy noise and something abnormal.

Dominant interference

Everything moves, so everything makes noise. On a bridge the traffic is the excitation — you are not detecting the vehicle, you are detecting the structure's response to it, which is a much smaller signal riding on a much larger one.

Mounting & protection

On the structure itself, never on the ground beside it. Bearings and expansion joints for a bridge, sleeper- or rail-adjacent for track, a cable clamp for the stay, a selected floor for a building. Mounting stiffness is part of the measurement, not an installation detail.

Array & sample rate

This group has the widest span in the article — 0.5 Hz – 20 kHz — so it genuinely needs both a dense and a sparse aperture on the same installation, with a peak requirement of 48 kHz.

Edge decision

Mixed by sub-case: pass-by is a seconds-scale judgement, rail impact is milliseconds, and cable and building trends are minutes. This is where a single fixed analysis path fails most visibly — which is why the group also shows every decision locus.

Output & response

Both local and cloud-reviewed decisions appear here. A single node can grade a pass-by on its own; a modal or tension change needs review across nodes or over time before anyone acts.

Basis & uncertainty

Pass-by speed from axle timing is straightforward physics. Structural modal shift as a damage indicator is [industry] practice, but the baseline for a specific bridge has to be recorded across a full temperature cycle before it means anything.

Scenarios in this group

ScenarioBand of interestSuggested sample rateJudgement time scaleDecision locus
Bridge — vehicle passage response, expansion-joint impact1 Hz – 2 kHz≥ 6 kHzsecondsDevice + cloud
Pavement — tyre noise as a proxy for damage and ponding200 Hz – 5 kHz≥ 12 kHzsecondsOn device
Rail track — wheel/rail noise, joints, corrugation, switches100 Hz – 20 kHz≥ 48 kHzmillisecondsOn device
Rope and stay cable — tension, wire break, rain-wind vibration0.5 Hz – 5 kHz≥ 12 kHzminutesTrigger + review
Building — curtain wall, structural noise, lifts, plumbing50 Hz – 10 kHz≥ 24 kHzsecondsOn device
Case 6 · Indoor & home

Indoors: Glass, Falls, Dry-Boil & Pipes

A room — a dwelling, a hotel room, a care facility. Quiet compared with every other case here, but the events are brief, the consequences of a miss are high, and there is a privacy constraint no other case has.

Site

Indoors, furnished, and quiet by industrial standards. Power and network are usually fine. The installation is consumer-facing, so appearance and discretion are functional requirements, not decoration.

What to listen for

Four distinct events: glass break (a sharp high-frequency crack followed by the tinkle of falling shards), a human fall (a low-frequency thump, then the absence of movement), a pot boiling dry, and a pressurised pipe leak (an ultrasonic hiss).

Dominant interference

The events are brief and the room is not empty. A door slam resembles a fall; a dropped plate resembles glass; running water masks a leak. Every one of these has to be rejected on features, not on level.

Mounting & protection

Ceiling-centre for a room, or integrated into a wall or ceiling panel. There is no cabinet and no trunking, so the enclosure has to look like part of the room.

Array & sample rate

The highest sample-rate requirement in the entire article appears here — 100 kHz — driven by the ultrasonic pipe-leak band up to 50 kHz. That single requirement pushes the front end far above what the other indoor events need, which is an argument for a separate, cheaper path for the low-rate events.

Edge decision

All four are local closed loop. Glass and fall are milliseconds-scale, dry-boil is seconds-scale, and a leak is a minutes-scale accumulation. Privacy makes local-only decisioning a design requirement rather than a preference.

Output & response

Local alert only. For a dwelling or a care room, no audio should leave the room at all — the output is the event, its class and a confidence. That is a product decision as much as an engineering one, and it should be made explicitly and early.

Basis & uncertainty

Glass-break and fall detection are well-published [industry] features. Dry-boil relies on a temperature-coupled acoustic change and is [engineering]. Any detection rate must be quoted for a specific room, microphone position and furnishing — we are not going to quote one here.

Scenarios in this group

ScenarioBand of interestSuggested sample rateJudgement time scaleDecision locus
Glass breakage2 – 8 kHz≥ 24 kHzmillisecondsOn device
Fall — impact plus vibration20 Hz – 2 kHz≥ 8 kHzmillisecondsOn device
Hob left dry-burning — whistle, vaporisation, empty pan1 – 16 kHz≥ 40 kHzsecondsOn device
Pipe — leak, cavitation, blockage100 Hz – 50 kHz≥ 100 kHzminutesOn device
Case 7 · Perimeter security

Airspace: Drone Intrusion

A perimeter — a site, a campus, a substation — where the airspace overhead is the thing being watched. Quiet, outdoor, and often the only sensor class that works at night without lighting anything up.

Site

An outdoor perimeter, usually quiet at night, with power and network at the mast. Multiple nodes are the norm, because one node gives you a bearing and not a location.

What to listen for

The harmonic stack of the rotor, and the Doppler shift as the source moves. Unlike every other case in this article, the target is the source's motion rather than its condition.

Dominant interference

Wind, foliage, road traffic, and above all your own site's HVAC. The drone is quiet and the sky is large, which makes this a detection-range problem long before it becomes a classification problem.

Mounting & protection

On a mast at the perimeter, as high and as clear as possible, weather-protected. Mast rigidity matters: a mast that moves in the wind modulates exactly the frequencies you are trying to track.

Array & sample rate

Group 1 scenario, band 100 Hz – 8 kHz, peak requirement 24 kHz. The aperture that gives useful bearing resolution is not the one that best serves the rotor harmonics — so the two apertures do genuinely different jobs here, rather than one serving as a spare.

Edge decision

A seconds-scale judgement, and local closed loop. A perimeter alert that waits for a cloud round trip is a perimeter alert that arrives late.

Output & response

Bearing, a harmonic-stack confidence, and a track over time. Audio retention is usually a policy question rather than a technical one, so treat raw capture as opt-in and document the choice.

Basis & uncertainty

Rotor harmonics are computable from blade count and RPM range, so the expected band is knowable in advance. Detection range against your own site's noise depends entirely on that site's background and must be measured — no vendor number transfers between sites.

Scenarios in this group

ScenarioBand of interestSuggested sample rateJudgement time scaleDecision locus
Drone intrusion — blade-pass frequency and harmonics100 Hz – 8 kHz≥ 24 kHzsecondsOn device
Where one window fails

Four Sites Need Two Windows, Not One

Ask each scenario two questions: what is the shortest window its lowest frequency can be resolved in, and how fast must the event be judged? In 4 of the 22 scenarios the analysis needs a longer window than the response time allows.

4Scenario / 22
50×Shortfall
6.6Minimum window
ScenarioAnalysis window ≥Event time scaleShortfall
Press / machine tool — impact, tool wear{sc_i-press_win}1 ms (milliseconds)5×
Rockfall / slope collapse — impact transient and rolling rhythm{sc_d-rock_win}1 ms (milliseconds)10×
Rail track — wheel/rail noise, joints, corrugation, switches{sc_t-rail_win}1 ms (milliseconds)10×
Fall — impact plus vibration{sc_h-fall_win}1 ms (milliseconds)50×

The answer is not a faster processor. It is two windows: a short trigger window that answers only is something happening, and a long analysis window that answers what is it and how bad. The trigger fires on a coarse feature; the analysis then runs behind it, on a buffer that already contains the beginning of the event. This is exactly what the hardware VAD's pre/post storage exists for — at the instant of a rockfall, the onset, which is the most diagnostic part of the whole record, is already stored.

And because the fastest case and the slowest case here differ by 6.6 orders of magnitude (3600000×), no single analysis path serves both ends. Several parallel paths, each with its own window and its own threshold, is the only structure that works — and it is a pipeline decision, not a tuning one.

From inference to commissioning

What Turns a Scenario Into an Installation

Everything above is design inference. Five steps stand between it and a number anyone can rely on. Two of them are site work, and no amount of engineering substitutes for them.

StepWhat happensWhat you get out of it
1 · SurveyRecord the site across a full operating cycle — shift changes, weather, the worst hourA real noise baseline and the actual interference inventory, instead of an assumed one
2 · Choose the aperture pairPick geometry from the band the site needs, not from a catalogueA sensor geometry — plus the knowledge of what it cannot hear
3 · Set the windowsShort trigger plus long analysis, separately per event classA response time and a resolution that are both defensible
4 · Calibrate thresholdsAgainst recorded events and recorded non-events from this siteFalse-alarm and miss rates you can actually quote to a customer
5 · Commission and drift-checkRe-baseline periodically as fans wear, buildings age and seasons turnA system that stays valid instead of decaying into nuisance alarms

Steps 1 and 4 are site work, not design work. Any figure quoted before them is an estimate, and we label it as one — including everything in this article.

Questions we get asked

FAQ

Can one RK3308 board cover all 22 scenarios?

The board can capture all of them — it has 8-channel PDM, two 8-channel I2S/TDM buses and 8 on-chip ADC channels. A single installation cannot, because the observation windows differ by 6.6 orders of magnitude. Multiple analysis paths on one board is the realistic answer.

Does RK3308 have an NPU?

No. It is 4×Cortex-A35 @ 1.3 GHz with no neural accelerator, so on-device inference means classical DSP plus small models. Anything heavier goes to the cloud or to a platform with an NPU — the same group's RK182X solution page covers that path.

Why not just stream the audio and decide in the cloud?

8 channels at 24 bit is already 93 GB per day per node at 1.10 MiB/s for a 192 kHz-class capture — before compression, retries or uplink cost. At the high-rate end it reaches 371 GB per day. The decision has to be made where the data is.

Is the pre-trigger buffer really necessary?

For every impulsive case, yes. The onset is the most diagnostic part of a rockfall, a glass break or a press stroke, and by the time a level threshold fires that onset is already in the past. The hardware VAD's pre/post storage is what makes the two-window structure possible at all.

How accurate are these systems?

We do not quote a figure, because for every case here accuracy depends on a baseline recorded at that specific site. What we can state is the method, the derived limits and the measurement plan. Anyone quoting a detection rate before site survey is quoting a guess.

Can the same hardware do condition monitoring and event detection?

Yes, and that is usually the sensible commercial answer — one front end, several analysis paths. What it cannot do is serve them with one window, so the paths must be designed as separate pipelines from the start.

Related

Continue exploring

Which of these sites is yours?

Tell us the site and the event you need to catch. We will say which of the 7 cases it resembles, what the aperture and window implications are, and what a survey would have to record before any threshold means anything.

Inquire All Tech Notes
—

Basis of figures, and what is not here

Frequency bands, aperture limits and data rates are derived here from first principles and carry the [physics] and [chip] labels. Judgement time scales and decision loci are design conclusions labelled [engineering]. Where a value is [industry] practice it is named as such. There is no measured field data in this article. No site in it has been instrumented by us, no client is referenced, and no accuracy figure is claimed. Section 5 exists precisely because the missing step — survey and threshold calibration on the real site — cannot be replaced by better analysis. Where a number must be measured to be meaningful, we have said so instead of supplying one.