Product: "DM2S/NX - The NextGen Document and Media Management System" (formerly DMS/NX, formerly ISE) Session: 2026-09-26 07:25 → 2026-09-27 05:40 AEST | Andrew.human & Claude Supersedes: DMS-NX__ISE_v9.6__Chapter4_Checkpoint_2026-09-26_2.md (kept, not overwritten). Chain: Chapter3_Checkpoint_2026-09-25_2 → Chapter4 → Chapter4_2 → this.
This is a full hand-over. §1 and §1b are what changed in Chapter 4 (§1b is new since _2). §4 onward carries forward everything from Chapter 3 that still applies.
Settings_ISE.json are not SMF code._2, _3 and so on.@version/@date/changelog), a visible runtime version, and a standalone Markdown manual.<placeholder>s.md5sum line per machine). Build a script only when Andrew asks for one. (Andrew: "u really complicate simple shit".)/bin on each machine (Andrew drives everything from there). ISE's own code stays in The_ISE_Project. /bin is not mirrored, so copies are updated by hand; the version printed on each run shows drift._meta verified after the run: revision 8fa0f1c6d0433df6e97c127f64b2a1d6c0dcda8a and model laion/clap-htsat-unfused. The overlap did no damage.9f6c8cf62f7204051cdd7677e5ef6f43 on both machines, 3,800,480 bytes. audio_checkpoint.json stays on Clone only (§7).transcribe_pieces_v1.0.0.zip, containing the script and transcribe_pieces_manual.md. It's installed at /root/transcribe_pieces.sh on Live.ffmpeg -f segment cuts it into 30 s pieces (no ffprobe needed), and llama-mtmd-cli runs Qwen3-ASR once per piece. It writes /root/transcripts/<name>.transcript.txt ([mm:ss] text, tag stripped) and .raw.txt.transcribe.sh:
Transcribe this audio., not the 7-word summary. It can be overridden with ASR_PROMPT.--temp 0 makes output repeatable.< /dev/null.LLAMA_BIN, ASR_MODEL and ASR_MMPROJ.KEEP=1 keeps the work folder.[!], and the run carried on.instagram-1790350777642.mp4 gave 3 pieces, 3 ok. The transcript is accurate ("What is a tensor? A language model works with lists of numbers…"). The text arrives on stdout. The runs took 761 s and 715 s./root/llama.cpp/build/bin/llama-mtmd-cli/root/models/qwen3-asr/Qwen3-ASR-0.6B-Q8_0.gguf (804,749,248 bytes)mmproj-Qwen3-ASR-0.6B-Q8_0.gguf (214,392,480 bytes)llama-server on either machine./rsync-live-root/ (a duplicate copy of the models, about 1.5 GB). Andrew's call._2 (2026-09-26 09:35 → 2026-09-27 05:40)mirrorISE.sh now has the Clone-only excludes, with its structure otherwise unchanged:
The_ISE_Project): AudioProcessor.py, AudioProcessor_manual.md, audio_vocab.txt, clap_diag.py, clap_smoketest.py.ISE_Data rsyncs: 'audio_store.json*' (also protects the .tmp) and audio_checkpoint.json./bin/json_push_audio_Clone_to_Live (v2.1.0, his rename of Claude's pushaudio_mirror.sh). It pushes audio_store.json to Live by IP, md5-checks both ends, refuses to push while AudioProcessor is running, then calls /bin/mirrorISE.sh.9f6c8cf6…, unchanged). The mirror pulled 4 files Live → Clone: ise_trace.log, misc_indexer.log, misc_store.json, word_index_misc.json. Only 13,004 bytes actually crossed the link for 172.6 MB of files (rsync delta); the 30 MB/s figures are local rebuild speed, not network speed. Nothing except audio_store.json is ever written to Live.-e "ssh -q" on the rsync lines (and ssh -q for the md5 call) would silence it..json FILES AGREE (2026-09-27 ~05:30)md5sum /var/www/html/ISE_Data/*.jsonaudio_checkpoint.json exists only on Clone, as designed..json files at the root of the main ISE_Data. smf20/smf21 are separate "proof of compatibility" installs with their own databases and ISE, so they're out of scope. Logs, traces and subfolders don't matter.audio_store.json 9f6c8cf62f7204051cdd7677e5ef6f43, misc_store.json 951dbcdcc0b56e689281125304fe739a, word_index_misc.json fba44685b7f1e6fc100197bd3665d775.json_md5_push.sh / json_md5_diff.sh v1.2.0 also exist (Clone+Dev push lists to /root/json_md5/ on Live; Live diffs). Andrew is keeping them in /bin, but the one-line md5sum is the primary method. json_md5_check.sh is superseded.spoken tag. Still undecided: opt-in by educational boards (named once, automatic) or by a hand-picked list of attachment IDs.*.mp3 returns nothing because MediaSearch 1.5.2 doesn't read audio_store.json yet (task 3). Wildcard support in MediaSearch is also unconfirmed.audio_store.json layout (confirmed on Live): {"_meta": {...}, "items": {"<id_attach>": {"filename", "extension", "duration", "class", "windows", "vector", ...}}}. Keys are SMF attachment IDs; the store does not record the post.eval $(php -r 'include "/var/www/html/Settings.php"; echo "DB=$db_name P=$db_prefix";')
mysql -p "$DB" -e "SELECT a.id_attach, a.filename, a.id_msg, m.id_topic, m.subject FROM ${P}attachments a LEFT JOIN ${P}messages m ON m.id_msg = a.id_msg WHERE a.id_attach IN (106, 108);"changelog.md and module doc-blocks instead of a full QCE chain. SMF native search remains his backup (finds what he wants ~99% of the time).Transcription speed diagnosis on Live. The work folder /tmp/transcribe.Aa1kQc is still there:
grep -iE 'eval time|tokens per second|n_threads|total time' /tmp/transcribe.Aa1kQc/piece_0000.log
wc -w /tmp/transcribe.Aa1kQc/piece_0000.out
nproc; free -m
Then run the same file on Clone for a direct comparison (pull it from Clone):
rsync --partial -a root@128.199.200.202:/root/instagram-1790350777642.mp4 /root/
KEEP=1 /root/transcribe_pieces.sh /root/instagram-1790350777642.mp4transcribe_pieces v1.0.1, after item 1:
-n cap if the model runs on after the transcriptllama-server rebuild so the model stays loadedaudio_store.json writer for opt-in educational items only (not every spoken item; see §1b transcription policy)What are 2959 (1h00m50s) and 6240 (1h01m31s)? Both are tagged male, lecture, interview, slow. If they're talks, they're the first real transcription targets.
AudioProcessor v1.0.4:
sec_per_window default 19.6 → 8.3Tag review (Chapter 3 task 2), now that full results exist. All of these can be fixed with a re-tag, no audio needed.
fast is on almost every item.slow, fast on 2846 and 2959, and dark, happy is common.country on dance tracks.music everywhere adds nothing.podcast/lecture false positives on music.--top-k 1, top-1 only for opposite groups (tempo, mood, vocals), and dropping groups that don't discriminate.Duplicate audio uploads. These have identical durations and tags:
The fix is the same SHA-256 "master copy" collapse planned for PDFs.
Chapter 3 tasks 3 and 5–15 carry forward unchanged:
MediaSearch.py uploaded; also check wildcard handling)The Chapter 2/3 transcript item is dropped: Andrew now makes the transcript PDFs himself.
Deploying MediaProcessor v9.21 on Live is still pending too.
Unresolved conflict: the Chapter 3 checkpoint says glibc 2.12 on Live, Clone and Dev. Earlier notes say Live and Clone were upgraded to 2.17. Run ldd --version | head -1 on each machine to settle it.
/root/transcribe.sh and /root/transcribe3.sh, and /var/log/commands confirms runs on 21, 22 and 24 September. .bash_history had no llama-mtmd-cli lines because the command was inside those scripts. The scripts were edited after their last run: transcribe3.sh on 25 September at 23:48 and transcribe.sh on 26 September at 01:56. The exact command is:
/root/llama.cpp/build/bin/llama-mtmd-cli \
-m /root/models/qwen3-asr/Qwen3-ASR-0.6B-Q8_0.gguf \
--mmproj /root/models/qwen3-asr/mmproj-Qwen3-ASR-0.6B-Q8_0.gguf \
--audio <16 kHz mono wav> \
-p "Transcribe this audio then summarize it in 7 words or less"Done. Dev's scripts were already complete, and Clone's mirrorISE.sh now has the Clone-only excludes (see §1b). Mirror runs on both machines are safe.
Low priority, both machines: in mirror1.sh, unanchored excludes (.htaccess, index.html, manifest.json, cache/) match at every depth, so new or changed files with those names deeper in the tree on Live (e.g. attachments/.htaccess) never reach the mirror. Existing copies are kept. To exclude only the top level, anchor the pattern with a leading slash, e.g. --exclude /.htaccess.
Live (forum, 128.199.200.202) |
Clone (clone, 192.168.233.140) |
Dev (dev, 192.168.233.129) |
|
|---|---|---|---|
| Arch | x86_64 | x86_64 | i686 (being retired) |
| RAM | 2 GB | 6 GB, 4 vCPU | n/a |
| Role | Production | 64-bit Dev; heavy AI jobs (audio, transcription) | Old Dev, keyword-only ISE |
| ffprobe | MISSING | /usr/local/bin/ffprobe |
/usr/bin/ffprobe (2.6.8) |
| tmux | yes | no (use nohup) |
n/a |
| File | Version | State |
|---|---|---|
MediaSearch.py |
1.5.2 | Dev and Clone. Live not updated (Andrew's call) |
MediaProcessor.py |
ISE v9.21 | Mock-tested, not deployed |
AudioProcessor.py |
1.0.3 | Clone only. Full run complete |
audio_vocab.txt |
1.0.0 | 104 labels, 6 groups, sha1 54504f3cce |
clap_diag.py |
1.1.0 | Clone |
dev_py_sync.sh |
1.0.2 | Dev |
transcribe_pieces.sh |
1.0.0 | Live. Clone is the target |
/bin/mirrorISE.sh (Clone) |
Andrew's, + Clone-only excludes | In use |
/bin/json_push_audio_Clone_to_Live |
2.1.0 (Andrew's naming) | Clone. Real run OK 2026-09-27 |
json_md5_push.sh / json_md5_diff.sh |
1.2.0 | /bin. Main ISE_Data only by default |
Live (images and video):
cd / && /usr/local/bin/python3.8 /var/www/html/The_ISE_Project/MediaIndexer.py
cd / && nohup /usr/local/bin/python3.8 /var/www/html/The_ISE_Project/MediaProcessor.py --log > /dev/null 2>&1 &
tail -f /var/www/html/ISE_Data/mediaprocessor.log
Clone (audio). Check that nothing is already running first:
pgrep -fl AudioProcessor
cd / && /root/venv-audio-test/bin/python /var/www/html/The_ISE_Project/AudioProcessor.py --plan
cd / && nohup /root/venv-audio-test/bin/python /var/www/html/The_ISE_Project/AudioProcessor.py --log --monitor > /dev/null 2>&1 &
tail -f /var/www/html/ISE_Data/audioprocessor.log
rsync --partial -a /var/www/html/ISE_Data/audio_store.json root@128.199.200.202:/var/www/html/ISE_Data/
Other operating notes:
pip install as root, run chmod -R go+rX /usr/local/lib/python3.8/site-packages.pgrep has no -a; use pgrep -fl, or ps -o pid,lstart,cmd.py_inventory.sh must run from / and must not be redirected./var/www/html/The_ISE_Project/ MediaSearch.py MediaProcessor.py MediaIndexer.py AudioProcessor.py
audio_vocab.txt clap_diag.py ise_settings.py qf_Mediasearch_bridge.php
/var/www/html/ISE_Data/ media_store.json embeddings_store.json captions_store.json
audio_store.json (Clone writes, pushed to Live) audio_checkpoint.json (Clone only)
audio_manifest.tsv hf_cache/ frame_cache/ (keep on Live) caption_tmp/
/root/transcribe_pieces.sh /root/transcripts/ /root/models/qwen3-asr/ /root/models/smolvlm/
/root/llama.cpp/build/bin/llama-mtmd-cli /root/venv-audio-test/ (Clone)
DM2S/NX Chapter 4 checkpoint _3, 2026-09-27 05:45 AEST