All 7 of these sites are different problems — and not because the microphone changes. They are different because how long you must listen changes. Each case below is written the same way: the site, the signature, the interference, the mounting, the array, the decision, the output, and honestly where the numbers come from.
How to read this. These are engineering scenarios, not delivered projects. No client is named, and none of these sites has been commissioned by us. Each case is a design inference built from four graded kinds of input, labelled throughout: [chip] Rockchip RK3308 product page and datasheet; [physics] derived here from the speed of sound and array geometry; [industry] published practice in vibration monitoring, acoustic emission and structural health monitoring; [engineering] our own design conclusions drawn from the three above. Where a value can only be meaningful after it is measured on site, we say so rather than invent one.
Frequency band, resolution and pipeline structure are set by three separate things. Confusing them is the most common way an acoustic monitoring project fails at the specification stage.
An 8-element array holds its directivity over only 1.8 octaves — and that span does not depend on the element spacing. Spacing slides the window; it never widens it. Widening needs more elements: 2.9 octaves at 16, 4.0 at 32.
Frequency resolution is Δf = 1/T, where T is the length of one analysis window. To resolve 0.1 Hz you must watch for at least 10 s; to resolve 20 kHz, 50 µs. That is the range a single 8-mic front end is being asked to span.
Across the 22 scenarios the judgement time scale runs from 1 ms to 3600 s — about 6.6 orders of magnitude (3600000×). It is this spread, not the band, that forces separate analysis paths.
Δf = 1/T is not a tuning choice, it is a floor: you cannot distinguish two features closer than 1/T apart. The right-hand column names the case in this article that sets each requirement.
| Frequency to resolve | Minimum window | Which case sets this |
|---|---|---|
| 0.1 Hz | 10 s | Tower modal drift — Case 2 |
| 1 Hz | 1 s | Bridge span response — Case 5 |
| 20 Hz | 50 ms | Fall impact, debris-flow onset — Cases 5, 4 |
| 100 Hz | 10 ms | Road and rail pass-by — Case 5 |
| 1 kHz | 1 ms | Dry-boil, fan blade pass — Cases 6, 3 |
| 8 kHz | 125 µs | Glass break — Case 6 |
| 20 kHz | 50 µs | Rockfall, gear mesh — Cases 4, 1 |
| 50 kHz | 20 µs | Pipe-leak ultrasound — Case 6 |
Each row is the shortest window in which that frequency is resolvable at all. In practice you take several windows' worth for a stable estimate — the numbers above are the floor, not the setting.
The consequence for this whole article: a dense and a sparse aperture can sit on the same board and serve two different frequency bands at once. But no arrangement of apertures makes the window flexible — window length is set by the event, and the events here differ by 6.6 orders of magnitude. Bands can share an installation. Windows cannot.
Each case is laid out identically so they can be compared side by side. The eight fields are the same eight questions you would ask a systems engineer in a kickoff meeting.
A press shop or a pump hall: motors, gearboxes and a press sharing one floor. Mains power, an existing cabinet, a plant network — and a background that is loud, broadband and shift-dependent.
Indoors, hard-surfaced, reverberant. Mains and network available, cabinet space exists, and maintenance staff already visit on a schedule. Background level is set by the machines themselves and changes with the shift.
Bearing defect tones and their sidebands against shaft harmonics, gear-mesh frequencies, and the impulsive signature of a press stroke. The signal of interest is rarely the loudest sound in the room — it is the part that changes.
The machine's own tonal content. An incipient bearing fault contributes far less energy than the running tones, so broadband level monitoring saturates long before the defect becomes visible. Narrow-band tracking of a specific order is what works.
On the machine frame, or on a mast beside it, protected to IP65 and rated for the hall's ambient temperature. Not on the floor slab if the press is on it — the slab carries the press impulse and will dominate everything.
The group spans 10 Hz – 20 kHz, with a peak requirement of 48 kHz. No single aperture covers that: expect a dense aperture for the 1 kHz-and-up gear and press content, plus a wider one for the low-order shaft tones. This is the two-aperture pattern from the Tech Note, in its simplest form.
Baseline per machine and per operating state, then track order-band energy against that baseline. Bearing degradation is an hours-scale judgement; a press stroke is a milliseconds-scale event. Same sensor, two very different windows — see the two-window section below.
Local closed loop. A graded alarm to the maintenance terminal plus a stored feature vector. No raw audio leaves the plant. Escalation to an engineer carries the trend, not the waveform.
Order and mesh frequencies come from geometry and speed, so they are computable from nameplate data. Bearing-defect band ratios are [industry] practice. Early-fault detectability and alarm thresholds are [engineering] and must be calibrated on the actual machine — we do not quote a universal threshold.
| Scenario | Band of interest | Suggested sample rate | Judgement time scale | Decision locus |
|---|---|---|---|---|
| Motor / pump — unbalance, misalignment, bearing | 10 Hz – 10 kHz | ≥ 24 kHz | hours | On device |
| Gearbox / reducer — mesh, tooth breakage | 100 Hz – 20 kHz | ≥ 48 kHz | hours | On device |
| Press / machine tool — impact, tool wear | 200 Hz – 20 kHz | ≥ 48 kHz | milliseconds | On device |
Three different structures with one thing in common: the interesting frequencies are extremely low, the sites are remote, and nobody is standing there.
An unattended outdoor site. A turbine tower, a transmission tower and a distribution transformer all present temperature extremes, continuous vibration and long maintenance intervals. Power is either available at the base or must be harvested.
Tower modal frequencies and their drift — a loss of stiffness shows up as a shift — plus blade-pass modulation, and transformer core and winding vibration at mains harmonics.
Wind. It is the loudest thing on site and it is not the signal. Gust loading, blade swish and rain all modulate the same low-frequency band you are trying to watch, and none of them is stationary.
Inside the tower base or on the nacelle; on the transformer tank wall; on the tower legs. The enclosure must survive weather without a maintenance visit, so corrosion and cable entry matter as much as the sensor.
The band starts at 0.1 Hz – 10 kHz — below the reach of any compact aperture, so a large sparse aperture is mandatory. This is also the group that needs two sample rates in parallel: a normal rate for the acoustic content, plus a slow channel for the sub-hertz trend.
Modal tracking is a minutes-to-hours judgement on a long window. This group is local-trigger-plus-cloud-review: a single node cannot separate a stiffness change from a wind event on its own.
Local trigger, plus a slow feature stream upward. At any useful sample rate the raw stream is enormous, so what leaves the site is modal frequencies, damping estimates and deviation trends — not audio.
Tower modal frequency depends on the real structure and its foundation, so it cannot be inferred from a datasheet and must be measured. Blade-pass and mains harmonics are computable from nameplate data. Damping estimation accuracy against a specific tower is [engineering].
| Scenario | Band of interest | Suggested sample rate | Judgement time scale | Decision locus |
|---|---|---|---|---|
| Wind turbine — main bearing, blade, yaw | 0.1 Hz – 10 kHz | ≥ 24 kHz + slow channel | hours | Trigger + review |
| Tower / blade — modal and low-frequency response | 0.1 – 20 Hz | ≥ 100 Hz | minutes | Trigger + review |
| Transformer / reactor — abnormal hum | 50 Hz – 5 kHz | ≥ 12 kHz | hours | On device |
A long, hard-surfaced shed: rows of ventilation fans at one end, animals in pens, an automatic feeding line — and a humidity, dust and wash-down regime that kills ordinary enclosures.
A harsh indoor environment by any measure: humid, dusty, ammonia-bearing air, regular high-pressure wash-down, and a long reverberant volume. The uplink is often unreliable, which pushes decisioning to the edge by default.
Two very different things. First, fan and drive degradation — bearing, belt and imbalance. Second, animal vocalisation: distress calls, coughing, and the shift in a group's sound when something in the house changes.
Reverberation. A hard shed with parallel walls is an acoustically hostile box; the direct sound from one pen is buried in the diffuse field. On top of that, the fans never stop, so there is no quiet reference to compare against.
Rigid mounting, positioned out of the fan airflow to avoid flow noise directly on the microphone, high enough to clear animals and wash-down spray. Sealing and cable routing matter more here than the sensor choice.
The group spans 20 Hz – 10 kHz with a peak requirement of 24 kHz. A mid aperture serves the fan content; animal vocalisation sits higher and needs less aperture but more channels, because separating pens is a spatial problem.
Fan degradation is a minutes-scale judgement against a per-fan baseline. Animal sound changes on a seconds scale. Both are local closed loop, because the shed may have no reliable uplink at all.
A graded fan-maintenance alarm and an animal-welfare indicator, both computed locally. What leaves the site is the indicator and its confidence — not the audio.
Fan defect frequencies are [industry] practice and need a per-model baseline. Animal vocalisation classifiers need species-specific data, so we describe the method and do not claim an accuracy figure.
| Scenario | Band of interest | Suggested sample rate | Judgement time scale | Decision locus |
|---|---|---|---|---|
| Ventilation fan — bearing, belt, rotor lock | 20 Hz – 10 kHz | ≥ 24 kHz | minutes | On device |
| Animal vocalisation — cough, alarm, farrowing | 50 Hz – 8 kHz | ≥ 24 kHz | seconds | On device |
| Feed line / scraper — jam and abnormal impact | 20 Hz – 8 kHz | ≥ 24 kHz | minutes | On device |
A mountain stream or a valley wall. No mains, no network, fully weather-exposed — and the event you care about may happen once a year, or once an hour.
A remote outdoor site with no infrastructure. Power is solar or harvested, the uplink is intermittent at best, and the unit must be assumed unreachable for months. Snow, frost and flood all have to be designed for.
A change in the river's own sound — a debris flow raises the bed-load impact rate and lowers the spectral centroid — and the impulsive crack-and-roll of a rockfall starting on a slope.
Rain and wind. Rain on the enclosure is broadband and impulsive, which makes it look a great deal like the events being hunted. This is the hardest false-positive problem in the whole set.
On a mast or a rock anchor, above flood level, aimed at the channel or the slope. The enclosure has to shed water and take an impact; the cabling has to survive frost and animal damage.
Flow-change judgement needs content across 10 Hz – 20 kHz, so wideband capture is warranted, and the rockfall is impulsive and needs the high end. Peak requirement 48 kHz.
The rockfall is a milliseconds-scale event — catch it with a short trigger window. The debris flow is a minutes-scale trend in continuous sound. Two windows at one site, which makes this group the clearest example of the two-stage structure.
Local decision plus cloud explanation. The local node triggers immediately; the cloud is used to separate a real event from rain by looking at longer context and at neighbouring nodes. A false alarm here costs a response team, so the two-stage split is a safety feature rather than an optimisation.
The acoustic signature of bed-load transport is [industry] research rather than a settled standard. Thresholds for a specific channel require recording across seasons; we do not offer a universal threshold.
| Scenario | Band of interest | Suggested sample rate | Judgement time scale | Decision locus |
|---|---|---|---|---|
| River acoustics — discharge and turbulence shift | 20 Hz – 5 kHz | ≥ 12 kHz | minutes | On device |
| Debris flow / flash flood — low rumble plus sustained scour | 10 Hz – 2 kHz | ≥ 6 kHz | minutes | Device + cloud |
| Rockfall / slope collapse — impact transient and rolling rhythm | 100 Hz – 20 kHz | ≥ 48 kHz | milliseconds | On device |
Five infrastructure types that share one driver: traffic. A bridge deck, a road surface, a rail track, a stay cable and a tall building — and in every one of them, the excitation is also the interference.
Outdoor, exposed, and in most cases impossible to close for installation. Access — not electronics — is the binding constraint on this group, and it shapes the whole deployment plan.
Vehicle pass-by signature (axle count and timing, hence speed), the structure's response to a known load, rail–wheel interaction, cable vibration and its damping, and in a building the difference between normal occupancy noise and something abnormal.
Everything moves, so everything makes noise. On a bridge the traffic is the excitation — you are not detecting the vehicle, you are detecting the structure's response to it, which is a much smaller signal riding on a much larger one.
On the structure itself, never on the ground beside it. Bearings and expansion joints for a bridge, sleeper- or rail-adjacent for track, a cable clamp for the stay, a selected floor for a building. Mounting stiffness is part of the measurement, not an installation detail.
This group has the widest span in the article — 0.5 Hz – 20 kHz — so it genuinely needs both a dense and a sparse aperture on the same installation, with a peak requirement of 48 kHz.
Mixed by sub-case: pass-by is a seconds-scale judgement, rail impact is milliseconds, and cable and building trends are minutes. This is where a single fixed analysis path fails most visibly — which is why the group also shows every decision locus.
Both local and cloud-reviewed decisions appear here. A single node can grade a pass-by on its own; a modal or tension change needs review across nodes or over time before anyone acts.
Pass-by speed from axle timing is straightforward physics. Structural modal shift as a damage indicator is [industry] practice, but the baseline for a specific bridge has to be recorded across a full temperature cycle before it means anything.
| Scenario | Band of interest | Suggested sample rate | Judgement time scale | Decision locus |
|---|---|---|---|---|
| Bridge — vehicle passage response, expansion-joint impact | 1 Hz – 2 kHz | ≥ 6 kHz | seconds | Device + cloud |
| Pavement — tyre noise as a proxy for damage and ponding | 200 Hz – 5 kHz | ≥ 12 kHz | seconds | On device |
| Rail track — wheel/rail noise, joints, corrugation, switches | 100 Hz – 20 kHz | ≥ 48 kHz | milliseconds | On device |
| Rope and stay cable — tension, wire break, rain-wind vibration | 0.5 Hz – 5 kHz | ≥ 12 kHz | minutes | Trigger + review |
| Building — curtain wall, structural noise, lifts, plumbing | 50 Hz – 10 kHz | ≥ 24 kHz | seconds | On device |
A room — a dwelling, a hotel room, a care facility. Quiet compared with every other case here, but the events are brief, the consequences of a miss are high, and there is a privacy constraint no other case has.
Indoors, furnished, and quiet by industrial standards. Power and network are usually fine. The installation is consumer-facing, so appearance and discretion are functional requirements, not decoration.
Four distinct events: glass break (a sharp high-frequency crack followed by the tinkle of falling shards), a human fall (a low-frequency thump, then the absence of movement), a pot boiling dry, and a pressurised pipe leak (an ultrasonic hiss).
The events are brief and the room is not empty. A door slam resembles a fall; a dropped plate resembles glass; running water masks a leak. Every one of these has to be rejected on features, not on level.
Ceiling-centre for a room, or integrated into a wall or ceiling panel. There is no cabinet and no trunking, so the enclosure has to look like part of the room.
The highest sample-rate requirement in the entire article appears here — 100 kHz — driven by the ultrasonic pipe-leak band up to 50 kHz. That single requirement pushes the front end far above what the other indoor events need, which is an argument for a separate, cheaper path for the low-rate events.
All four are local closed loop. Glass and fall are milliseconds-scale, dry-boil is seconds-scale, and a leak is a minutes-scale accumulation. Privacy makes local-only decisioning a design requirement rather than a preference.
Local alert only. For a dwelling or a care room, no audio should leave the room at all — the output is the event, its class and a confidence. That is a product decision as much as an engineering one, and it should be made explicitly and early.
Glass-break and fall detection are well-published [industry] features. Dry-boil relies on a temperature-coupled acoustic change and is [engineering]. Any detection rate must be quoted for a specific room, microphone position and furnishing — we are not going to quote one here.
| Scenario | Band of interest | Suggested sample rate | Judgement time scale | Decision locus |
|---|---|---|---|---|
| Glass breakage | 2 – 8 kHz | ≥ 24 kHz | milliseconds | On device |
| Fall — impact plus vibration | 20 Hz – 2 kHz | ≥ 8 kHz | milliseconds | On device |
| Hob left dry-burning — whistle, vaporisation, empty pan | 1 – 16 kHz | ≥ 40 kHz | seconds | On device |
| Pipe — leak, cavitation, blockage | 100 Hz – 50 kHz | ≥ 100 kHz | minutes | On device |
A perimeter — a site, a campus, a substation — where the airspace overhead is the thing being watched. Quiet, outdoor, and often the only sensor class that works at night without lighting anything up.
An outdoor perimeter, usually quiet at night, with power and network at the mast. Multiple nodes are the norm, because one node gives you a bearing and not a location.
The harmonic stack of the rotor, and the Doppler shift as the source moves. Unlike every other case in this article, the target is the source's motion rather than its condition.
Wind, foliage, road traffic, and above all your own site's HVAC. The drone is quiet and the sky is large, which makes this a detection-range problem long before it becomes a classification problem.
On a mast at the perimeter, as high and as clear as possible, weather-protected. Mast rigidity matters: a mast that moves in the wind modulates exactly the frequencies you are trying to track.
Group 1 scenario, band 100 Hz – 8 kHz, peak requirement 24 kHz. The aperture that gives useful bearing resolution is not the one that best serves the rotor harmonics — so the two apertures do genuinely different jobs here, rather than one serving as a spare.
A seconds-scale judgement, and local closed loop. A perimeter alert that waits for a cloud round trip is a perimeter alert that arrives late.
Bearing, a harmonic-stack confidence, and a track over time. Audio retention is usually a policy question rather than a technical one, so treat raw capture as opt-in and document the choice.
Rotor harmonics are computable from blade count and RPM range, so the expected band is knowable in advance. Detection range against your own site's noise depends entirely on that site's background and must be measured — no vendor number transfers between sites.
| Scenario | Band of interest | Suggested sample rate | Judgement time scale | Decision locus |
|---|---|---|---|---|
| Drone intrusion — blade-pass frequency and harmonics | 100 Hz – 8 kHz | ≥ 24 kHz | seconds | On device |
Ask each scenario two questions: what is the shortest window its lowest frequency can be resolved in, and how fast must the event be judged? In 4 of the 22 scenarios the analysis needs a longer window than the response time allows.
| Scenario | Analysis window ≥ | Event time scale | Shortfall |
|---|---|---|---|
| Press / machine tool — impact, tool wear | {sc_i-press_win} | 1 ms (milliseconds) | 5× |
| Rockfall / slope collapse — impact transient and rolling rhythm | {sc_d-rock_win} | 1 ms (milliseconds) | 10× |
| Rail track — wheel/rail noise, joints, corrugation, switches | {sc_t-rail_win} | 1 ms (milliseconds) | 10× |
| Fall — impact plus vibration | {sc_h-fall_win} | 1 ms (milliseconds) | 50× |
The answer is not a faster processor. It is two windows: a short trigger window that answers only is something happening, and a long analysis window that answers what is it and how bad. The trigger fires on a coarse feature; the analysis then runs behind it, on a buffer that already contains the beginning of the event. This is exactly what the hardware VAD's pre/post storage exists for — at the instant of a rockfall, the onset, which is the most diagnostic part of the whole record, is already stored.
And because the fastest case and the slowest case here differ by 6.6 orders of magnitude (3600000×), no single analysis path serves both ends. Several parallel paths, each with its own window and its own threshold, is the only structure that works — and it is a pipeline decision, not a tuning one.
Everything above is design inference. Five steps stand between it and a number anyone can rely on. Two of them are site work, and no amount of engineering substitutes for them.
| Step | What happens | What you get out of it |
|---|---|---|
| 1 · Survey | Record the site across a full operating cycle — shift changes, weather, the worst hour | A real noise baseline and the actual interference inventory, instead of an assumed one |
| 2 · Choose the aperture pair | Pick geometry from the band the site needs, not from a catalogue | A sensor geometry — plus the knowledge of what it cannot hear |
| 3 · Set the windows | Short trigger plus long analysis, separately per event class | A response time and a resolution that are both defensible |
| 4 · Calibrate thresholds | Against recorded events and recorded non-events from this site | False-alarm and miss rates you can actually quote to a customer |
| 5 · Commission and drift-check | Re-baseline periodically as fans wear, buildings age and seasons turn | A system that stays valid instead of decaying into nuisance alarms |
Steps 1 and 4 are site work, not design work. Any figure quoted before them is an estimate, and we label it as one — including everything in this article.
The board can capture all of them — it has 8-channel PDM, two 8-channel I2S/TDM buses and 8 on-chip ADC channels. A single installation cannot, because the observation windows differ by 6.6 orders of magnitude. Multiple analysis paths on one board is the realistic answer.
No. It is 4×Cortex-A35 @ 1.3 GHz with no neural accelerator, so on-device inference means classical DSP plus small models. Anything heavier goes to the cloud or to a platform with an NPU — the same group's RK182X solution page covers that path.
8 channels at 24 bit is already 93 GB per day per node at 1.10 MiB/s for a 192 kHz-class capture — before compression, retries or uplink cost. At the high-rate end it reaches 371 GB per day. The decision has to be made where the data is.
For every impulsive case, yes. The onset is the most diagnostic part of a rockfall, a glass break or a press stroke, and by the time a level threshold fires that onset is already in the past. The hardware VAD's pre/post storage is what makes the two-window structure possible at all.
We do not quote a figure, because for every case here accuracy depends on a baseline recorded at that specific site. What we can state is the method, the derived limits and the measurement plan. Anyone quoting a detection rate before site survey is quoting a guess.
Yes, and that is usually the sensible commercial answer — one front end, several analysis paths. What it cannot do is serve them with one window, so the paths must be designed as separate pipelines from the start.
The engineering note behind this article: aperture geometry, the 1.8-octave invariant, sample-rate and data-rate derivations.
Solution · ENThe deliverable form of the same front end on metoclaw: what is delivered, integration surface and boundaries.
Case · ENWhat to get right first when the array's job is speech rather than diagnosis — microphone, sealing and placement.
Tell us the site and the event you need to catch. We will say which of the 7 cases it resembles, what the aperture and window implications are, and what a survey would have to record before any threshold means anything.
Inquire All Tech NotesFrequency bands, aperture limits and data rates are derived here from first principles and carry the [physics] and [chip] labels. Judgement time scales and decision loci are design conclusions labelled [engineering]. Where a value is [industry] practice it is named as such. There is no measured field data in this article. No site in it has been instrumented by us, no client is referenced, and no accuracy figure is claimed. Section 5 exists precisely because the missing step — survey and threshold calibration on the real site — cannot be replaced by better analysis. Where a number must be measured to be meaningful, we have said so instead of supplying one.