@version 1.0.3 · @date 2026-09-26 · Another collaboration between Andrew.human and Claude.ai
Part of DMS/NX (ISEmedia). This manual is written to be read by people and by AI assistants.
AudioProcessor turns the forum's audio attachments (mp3, wav, flac, ogg, m4a, aac, wma) into searchable tags, using the CLAP model laion/clap-htsat-unfused. The tags are stored as plain text in ISE_Data/audio_store.json, and MediaSearch on Live matches search words against that text. Live never runs CLAP. CLAP needs about 1.6 GB of memory, so AudioProcessor runs on Clone, and only the finished store is copied to Live.
| File | Where | Role |
|---|---|---|
AudioProcessor.py |
/var/www/html/The_ISE_Project/ (next to MediaProcessor.py) |
The program |
audio_vocab.txt |
same folder | The tag vocabulary. Edit freely |
ISE_Data/audio_store.json |
written by the program | Vectors, tags and captions. Clone is its only writer |
ISE_Data/audio_checkpoint.json |
written by the program | Human-readable progress |
ISE_Data/audioprocessor.log |
with --log |
Append-only run history |
ISE_Data/hf_cache/models--laion--clap-htsat-unfused |
pre-downloaded | The model, pinned to snapshot 8fa0f1c6… |
It depends on Collabware/core_utils.py (v1.2.0) to find attachment files, on ISE_Data/media_store.json (from MediaIndexer.py) for the list of audio items, and on ffmpeg and ffprobe being on the PATH.
Always run it with the audio venv's Python, from /:
cd / && /root/venv-audio-test/bin/python /var/www/html/The_ISE_Project/AudioProcessor.py --plan
--plan finds and measures every pending item, prints the number of windows and the estimated time, and encodes nothing.
To run it for real in the background (Clone has no tmux):
cd / && nohup /root/venv-audio-test/bin/python /var/www/html/The_ISE_Project/AudioProcessor.py --log --monitor \
> /dev/null 2>&1 &
tail -f /var/www/html/ISE_Data/audioprocessor.log
With no arguments it does a real run and picks up where the last one stopped (ISE CLI convention).
| Class | Rule | Stored |
|---|---|---|
| Regular (≤ 15 min) | 6 evenly spaced 10-second windows. Shorter tracks get fewer: one per 10 seconds of audio | One whole-track vector, its tags and caption |
| Long mix (> 15 min) | A 10-second window starting every 30 seconds | Every window as <id>_a<start_sec> with its own tags, plus the whole-track average |
Window keys never contain _t. That suffix belongs to video frames, and MediaSearch relies on it to recognise them.
Changing --long-threshold, --stride or --samples once a store exists is refused unless you add --force, which rebuilds everything. A store never mixes two sampling schemes.
Each vector is compared with every label in audio_vocab.txt. Within each [group], the labels compete with each other: the best one is always kept, plus up to --top-k (default 2) that score at least half as well. The chosen keywords become the item's caption, e.g. techno, dark, fast, instrumental, synth, music.
Changing the vocabulary is cheap. The 512-number vectors are saved. When the vocabulary file changes, the next run re-tags every stored item from those saved vectors in seconds, and no audio is processed again.
Vocabulary format:
[genre] {} music <- group header; {} is replaced by the keyword
techno <- CLAP hears "techno music"
hi-nrg | hi-nrg dance music <- CLAP hears exactly the text after |
# comment
| Option | Default | Meaning |
|---|---|---|
--plan |
off | Measure and estimate only |
--attachid N |
n/a | Only this attachment |
--chunk-size N |
all | Stop after N items |
--force |
off | Redo items already stored. If the settings changed, rebuild everything |
--threads N |
all cores | torch CPU threads |
--batch N |
4 | Windows per CLAP call |
--long-threshold S |
900 | Seconds; anything longer is a long mix |
--stride S |
30 | Seconds between long-mix windows |
--samples N |
6 | Windows per regular track |
--top-k N |
2 | Maximum tags per group |
--vocab PATH |
audio_vocab.txt next to the script |
A different vocabulary file |
--sec-per-window S |
19.6 | Rate used for --plan's time estimate |
--log [PATH] |
off | Also append output to a log file |
--any-host |
off | Run on a host other than Clone. Normally refused |
--monitor, --verbose |
off | Telemetry at the end; per-item plan and file-lookup detail |
From Clone:
rsync --partial -a /var/www/html/ISE_Data/audio_store.json root@128.199.200.202:/var/www/html/ISE_Data/
Clone's mirror from Live must exclude ISE_Data/audio_store.json* and ISE_Data/audio_checkpoint.json. Otherwise a mirror run deletes or overwrites Clone's copy.
8fa0f1c6…). A store started with v1.0.2 carries straight on, with nothing rebuilt.laion/clap-htsat-unfused. larger_clap_music gave every sound nearly the same vector under transformers 4.27.4 (found with clap_diag.py), so every track got identical tags. Also silences the tokenizers warning. A store built with the old model must be rebuilt.clone, because the first real run was attempted on Live. --any-host overrides.clap_smoketest.py v0.1.0, which is proven on Clone. Python 3.8-compatible (checked with vermin).Live has glibc 2.12, so the Rust tokenizers package can't be installed there. From v1.0.4, AudioProcessor falls back to the pure-Python RoBERTa tokenizer when tokenizers is missing. It produces the same token ids, so the tags match Clone's.
One-time setup on Live (Python 3.8):
/usr/local/bin/python3.8 -m pip install --only-binary=:all: "huggingface-hub>=0.11,<0.17" regex requests pyyaml filelock tqdm packaging
/usr/local/bin/python3.8 -m pip install --no-deps "transformers==4.27.4"
mkdir -p /root/shim && cat > /root/shim/ffprobe <<'SH'
#!/bin/sh
# Stand-in for ffprobe's duration query only (AudioProcessor's probe_duration).
for f; do :; done
ffmpeg -nostdin -i "$f" 2>&1 | awk '/Duration:/{split($2,t,":"); sub(",","",t[3]); print t[1]*3600+t[2]*60+t[3]; exit}'
SH
chmod 755 /root/shim/ffprobe
To classify one file:
ID=6500
cd / && /usr/local/bin/python3.8 /var/www/html/The_ISE_Project/MediaIndexer.py --attachid $ID
cd / && PATH=/root/shim:$PATH /usr/local/bin/python3.8 /var/www/html/The_ISE_Project/AudioProcessor.py --attachid $ID --any-host --threads 1 --log
Afterwards, and before Clone next runs or pushes: pull Live's audio_store.json back to Clone, or the next Clone-to-Live push will overwrite this item:
rsync --partial -a root@128.199.200.202:/var/www/html/ISE_Data/audio_store.json /var/www/html/ISE_Data/
These are measured by audio_features.py, which needs only numpy: no new packages and no new model. It measures the same decoded 10-second windows CLAP uses, at about 0.1 s per window.
| Where | Field | Meaning |
|---|---|---|
| each window (long mixes) | bpm, bpm_conf |
tempo, folded into 70–180. Below 0.2 confidence it's stored as null (no clear beat) |
key, key_conf |
e.g. A minor. Below 0.5 confidence it's null |
|
level_db |
RMS level, dBFS | |
| each item | bpm, bpm_range |
median, and the 10th–90th percentile |
key |
from the whole item's pitch profile | |
lufs, lra |
ffmpeg ebur128 over the whole file | |
changes |
long mixes: window starts where at least two of tempo, key and tags change (an approximate tracklist) |
New runs measure automatically. Existing items need one backfill run, on Clone. It doesn't load CLAP and doesn't touch vectors or tags:
cd / && /root/venv-audio-test/bin/python /var/www/html/The_ISE_Project/AudioProcessor.py --features-only --log
--force with --features-only re-measures everything. --attachid N limits it to one item. Then push audio_store.json to Live as usual.
Self-test of the measurements, which runs anywhere:
python3 /var/www/html/The_ISE_Project/audio_features.py --selftest
--verbose shows what the run is doing and everything it detected, item by item. It works in normal runs and with --features-only:
cd / && /root/venv-audio-test/bin/python /var/www/html/The_ISE_Project/AudioProcessor.py --attachid 6982 --force --verbose --log
For each item it prints:
Without --verbose, the output is the same as 1.1.0. The attachment-lookup debug that --verbose used to switch on is now --debug.