📝 AudioProcessor_manual.mdv4.5.1 · 2026-10-02

AudioProcessor.py: Manual

@version 1.0.3 · @date 2026-09-26 · Another collaboration between Andrew.human and Claude.ai

Part of DMS/NX (ISEmedia). This manual is written to be read by people and by AI assistants.


1. What it is

AudioProcessor turns the forum's audio attachments (mp3, wav, flac, ogg, m4a, aac, wma) into searchable tags, using the CLAP model laion/clap-htsat-unfused. The tags are stored as plain text in ISE_Data/audio_store.json, and MediaSearch on Live matches search words against that text. Live never runs CLAP. CLAP needs about 1.6 GB of memory, so AudioProcessor runs on Clone, and only the finished store is copied to Live.

2. Files

File Where Role
AudioProcessor.py /var/www/html/The_ISE_Project/ (next to MediaProcessor.py) The program
audio_vocab.txt same folder The tag vocabulary. Edit freely
ISE_Data/audio_store.json written by the program Vectors, tags and captions. Clone is its only writer
ISE_Data/audio_checkpoint.json written by the program Human-readable progress
ISE_Data/audioprocessor.log with --log Append-only run history
ISE_Data/hf_cache/models--laion--clap-htsat-unfused pre-downloaded The model, pinned to snapshot 8fa0f1c6…

It depends on Collabware/core_utils.py (v1.2.0) to find attachment files, on ISE_Data/media_store.json (from MediaIndexer.py) for the list of audio items, and on ffmpeg and ffprobe being on the PATH.

3. How to run it (on Clone)

Always run it with the audio venv's Python, from /:

cd / && /root/venv-audio-test/bin/python /var/www/html/The_ISE_Project/AudioProcessor.py --plan

--plan finds and measures every pending item, prints the number of windows and the estimated time, and encodes nothing.

To run it for real in the background (Clone has no tmux):

cd / && nohup /root/venv-audio-test/bin/python /var/www/html/The_ISE_Project/AudioProcessor.py --log --monitor \
> /dev/null 2>&1 &
tail -f /var/www/html/ISE_Data/audioprocessor.log

With no arguments it does a real run and picks up where the last one stopped (ISE CLI convention).

4. How audio is sampled

Class Rule Stored
Regular (≤ 15 min) 6 evenly spaced 10-second windows. Shorter tracks get fewer: one per 10 seconds of audio One whole-track vector, its tags and caption
Long mix (> 15 min) A 10-second window starting every 30 seconds Every window as <id>_a<start_sec> with its own tags, plus the whole-track average

Window keys never contain _t. That suffix belongs to video frames, and MediaSearch relies on it to recognise them.

Changing --long-threshold, --stride or --samples once a store exists is refused unless you add --force, which rebuilds everything. A store never mixes two sampling schemes.

5. How tagging works

Each vector is compared with every label in audio_vocab.txt. Within each [group], the labels compete with each other: the best one is always kept, plus up to --top-k (default 2) that score at least half as well. The chosen keywords become the item's caption, e.g. techno, dark, fast, instrumental, synth, music.

Changing the vocabulary is cheap. The 512-number vectors are saved. When the vocabulary file changes, the next run re-tags every stored item from those saved vectors in seconds, and no audio is processed again.

Vocabulary format:

[genre] {} music          <- group header; {} is replaced by the keyword
techno <- CLAP hears "techno music"
hi-nrg | hi-nrg dance music <- CLAP hears exactly the text after |
# comment

6. Options

Option Default Meaning
--plan off Measure and estimate only
--attachid N n/a Only this attachment
--chunk-size N all Stop after N items
--force off Redo items already stored. If the settings changed, rebuild everything
--threads N all cores torch CPU threads
--batch N 4 Windows per CLAP call
--long-threshold S 900 Seconds; anything longer is a long mix
--stride S 30 Seconds between long-mix windows
--samples N 6 Windows per regular track
--top-k N 2 Maximum tags per group
--vocab PATH audio_vocab.txt next to the script A different vocabulary file
--sec-per-window S 19.6 Rate used for --plan's time estimate
--log [PATH] off Also append output to a log file
--any-host off Run on a host other than Clone. Normally refused
--monitor, --verbose off Telemetry at the end; per-item plan and file-lookup detail

7. Safety and errors

8. After a run: publish to Live

From Clone:

rsync --partial -a /var/www/html/ISE_Data/audio_store.json root@128.199.200.202:/var/www/html/ISE_Data/

Clone's mirror from Live must exclude ISE_Data/audio_store.json* and ISE_Data/audio_checkpoint.json. Otherwise a mirror run deletes or overwrites Clone's copy.

9. Version history

Running a single file on Live (v1.0.4)

Live has glibc 2.12, so the Rust tokenizers package can't be installed there. From v1.0.4, AudioProcessor falls back to the pure-Python RoBERTa tokenizer when tokenizers is missing. It produces the same token ids, so the tags match Clone's.

One-time setup on Live (Python 3.8):

/usr/local/bin/python3.8 -m pip install --only-binary=:all: "huggingface-hub>=0.11,<0.17" regex requests pyyaml filelock tqdm packaging
/usr/local/bin/python3.8 -m pip install --no-deps "transformers==4.27.4"
mkdir -p /root/shim && cat > /root/shim/ffprobe <<'SH'
#!/bin/sh
# Stand-in for ffprobe's duration query only (AudioProcessor's probe_duration).
for f; do :; done
ffmpeg -nostdin -i "$f" 2>&1 | awk '/Duration:/{split($2,t,":"); sub(",","",t[3]); print t[1]*3600+t[2]*60+t[3]; exit}'
SH
chmod 755 /root/shim/ffprobe

To classify one file:

ID=6500
cd / && /usr/local/bin/python3.8 /var/www/html/The_ISE_Project/MediaIndexer.py --attachid $ID
cd / && PATH=/root/shim:$PATH /usr/local/bin/python3.8 /var/www/html/The_ISE_Project/AudioProcessor.py --attachid $ID --any-host --threads 1 --log

Afterwards, and before Clone next runs or pushes: pull Live's audio_store.json back to Clone, or the next Clone-to-Live push will overwrite this item:

rsync --partial -a root@128.199.200.202:/var/www/html/ISE_Data/audio_store.json /var/www/html/ISE_Data/

v1.1.0: BPM, key, loudness, track changes

These are measured by audio_features.py, which needs only numpy: no new packages and no new model. It measures the same decoded 10-second windows CLAP uses, at about 0.1 s per window.

Where Field Meaning
each window (long mixes) bpm, bpm_conf tempo, folded into 70–180. Below 0.2 confidence it's stored as null (no clear beat)
key, key_conf e.g. A minor. Below 0.5 confidence it's null
level_db RMS level, dBFS
each item bpm, bpm_range median, and the 10th–90th percentile
key from the whole item's pitch profile
lufs, lra ffmpeg ebur128 over the whole file
changes long mixes: window starts where at least two of tempo, key and tags change (an approximate tracklist)

New runs measure automatically. Existing items need one backfill run, on Clone. It doesn't load CLAP and doesn't touch vectors or tags:

cd / && /root/venv-audio-test/bin/python /var/www/html/The_ISE_Project/AudioProcessor.py --features-only --log

--force with --features-only re-measures everything. --attachid N limits it to one item. Then push audio_store.json to Live as usual.

Self-test of the measurements, which runs anywhere:

python3 /var/www/html/The_ISE_Project/audio_features.py --selftest

v1.1.1: --verbose

--verbose shows what the run is doing and everything it detected, item by item. It works in normal runs and with --features-only:

cd / && /root/venv-audio-test/bin/python /var/www/html/The_ISE_Project/AudioProcessor.py --attachid 6982 --force --verbose --log

For each item it prints:

  1. When it starts: the ID, file name, size, path, duration, class (regular or long) and how many windows it will take.
  2. While CLAP encodes: a progress line after every batch, with windows done so far and the seconds per window.
  3. When the item is stored:
    • tempo (median and range)
    • key, with its confidence
    • RMS level, LUFS and LRA
    • for long mixes, the approximate track changes as h:mm:ss
    • every tag with its probability, grouped as genre, mood, energy, vocals, instrument and sound
    • a table with one line per window: time, BPM and confidence, key and confidence, level in dB, and the window's top 5 tags (long mixes only; regular windows aren't tagged)

Without --verbose, the output is the same as 1.1.0. The attachment-lookup debug that --verbose used to switch on is now --debug.