dopetalk does not endorse any advertised product nor does it accept any liability for it's use or misuse


Our Discord Notification Server invitation link is https://discord.gg/jB2qmRrxyD

Author Topic: Master Post 01: "DM2S/NX - The NextGen Document and Media Mangagment System"  (Read 131 times)

Offline Chip (OP)

  • Server Admin
  • Admin
  • Hero Member
  • ****
  • Administrator
  • *****
  • Restrictive
  • *
  • RC Discussion
  • **
  • Private
  • ***
  • Owner
  • *****
  • Join Date: Dec 2014
  • Location: Australia
  • Posts: 7398
  • Reputation Power: 0
  • Chip has hidden their reputation power
  • Gender: Male
  • Last Login:Today at 06:36:28 AM
  • Deeply Confused Learner
  • Profession: IT Engineer now retired
πŸ“Ž Click on me to view all my attachments
ListAttBBC  |  v6.12  |  2026-09-10  |  Andrew.human & Claude.ai
#FileSizeDownloadsInfoDL
00audioprocessor.log14.7 KB3ℹ️⬇️
01Lou Reed's _Walk on the Wildside_TO_Hopw_to_Download_youtube_Videos_as_mp3_or_mp4.pdf  [native]1.1 MB6ℹ️⬇️
02Horalexia_ decoding a linguistic anomaly - Claude.pdf  [native]288.1 KB5ℹ️⬇️
03DMS-NX__ISE_v9.7__Chapter4_Checkpoint_2026-09-26_3.md17.7 KB2ℹ️⬇️
04DM2S_NX next generation document and media management - Chapter 4 - FINAL.pdf  [native]4.3 MB6ℹ️⬇️
05ISE_v9.7.zip7 MB6ℹ️⬇️
06DM2S_NX - The Next Generation Document and Media Management System v9.8 - Chapter 5 - Claude.pdf  [native]859.9 KB3ℹ️⬇️
07DM2S-NX_QAT_softcopy_v1.1.0.zip32.9 KB1ℹ️⬇️
08ISE_v9.8.zip7.3 MB3ℹ️⬇️
09DM2S_NX - The Next Generation Document and Media Management System v9.8 - Chapter 5 (CLOSED) - Claude.pdf  [native]4.6 MB5ℹ️⬇️
10all_attached_audio_features.txt123.1 KB18ℹ️⬇️
11DM2S-NX_v9.9_Chapter6_Checkpoint_2026-09-28.md4.7 KB1ℹ️⬇️
12Audio file download instead of playback - Claude.pdf  [native]313.9 KB1ℹ️⬇️
13yt-dlp SoundCloud download - Claude.pdf  [native]426.4 KB0ℹ️⬇️
14ISE search sorting and display issues - Claude.pdf  [native]708.2 KB0ℹ️⬇️
15all_attached_audio_features.txt123.1 KB3ℹ️⬇️
16ISE_Python_Scripts_Manual.md50.7 KB0ℹ️⬇️
17Win-to-Linux-traps-for-young-players_NOTES_readme.txt607 B2ℹ️⬇️
18PDFInder.py MUST BE RUN FROM Python 3.6 ONLY.pdf  [native]357.9 KB0ℹ️⬇️
19DM2S_NX - The Next Generation Document and Media Management System v9.9 - Chapter 6 - PARTIAL.pdf  [native]3.3 MB0ℹ️⬇️
20ISE_v9.9.zip7 MB0ℹ️⬇️
Master Post 01:  "DM2S/NX - The NextGen Document and Media Mangagment System"

TO-DO:

1. Classify/Index a single item

2. Log a summary of the search RESULTS alond with the search REQUEST > "searches"

3. Implement incremental DB upates for Dev and Clone for MIRRORING from Live,  but still copy in full

etc.

9. Incorporate ISEmedia with ISEmega.

With a new name, this continues on from:
* "Master Post 00: 'DMS/NX - The NextGen Document Mangagment System'"

Includes these SMF mods: Collaborative, ListAtt, PDFBBC, MDBBC, LMVBBC, LTVBBC and (the) ISE (Project) [stable at version 9.7]

Related aliases in "/.bashrc":

alias py='echo Use py36 or py38'
alias py36='python3.6'
alias py38='python3.8'
alias py3='python3.8'
alias py38='python3.8'
alias pip='pip3.8'
alias pip38='pip3.8'

πŸ“ ISE_Python_Scripts_Manual.mdv4.5 · 2026-09-28

ISE Project β€” Python Scripts Manual

Covers every Python script currently in /var/www/html/The_ISE_Project/. There are now five search sources (ISE/posts, ISEpdf, ISEmisc, ISEmega, ISEmedia) built on two different foundations β€” worth knowing before debugging any of them:

  • search.py / PDFsearch.py / MiscSearch.py / MegaSearch.py / MediaSearch.py (the five search scripts, plus IndexBuilder.py/ PDFIndexer.py/MiscIndexer.py) auto-detect DB credentials and the board URL via ise_settings.py's get_base_url()/get_project_url() β€” which as of v4.0/4.1 reads a small per-install Settings_ISE.json ({"environment": ..., "install": ...}) and looks the real URL up in a hand-verified table, rather than scraping/guessing it out of Settings.php. This corrects an earlier version of this manual, which described DB/URL config as coming "from Settings.php via ISE_ROOT" β€” that was true before ise_settings.py v4.0 (2026-08-21) replaced it; ISE_ROOT is still set by some bridges as an environment variable but nothing in current Python actually reads it back out. pdf_duplicate_spotter.py reuses PDFIndexer.py/PDFsearch.py's existing infrastructure directly rather than duplicating it.
  • MediaIndexer.py and MediaProcessor.py (the two scripts that build ISEmedia's index β€” not MediaSearch.py itself, see below) come from a separate codebase, "Collabware" (credited in their own docblocks as a 4-way effort: Andrew + Claude + ChatGPT + Gemini), and load settings through Collabware.core_utils (get_board_root()/get_data_dir()/ load_smf_settings()), not ise_settings.py. They also carry their own version-stamp convention β€” a global-sounding tag like "ISE v9.7"/ "ISE v9.18" in the docblock plus a RUNTIME_VERSION string kept manually in sync with it β€” which is not the same convention as every other ISE module's own per-file semver (search.py v7.5.2, PDFsearch.py's own version, etc.). Not a bug β€” just a real seam between two codebases glued together under one project, worth knowing so a version number or a settings-lookup error in one of these two scripts isn't chased through ise_settings.py/Settings_ISE.json by mistake. MediaSearch.py (the actual search script, unlike the two indexers) does use ise_settings.py like the rest of ISE β€” see its own section below.

PDFIndexer.py

Builds the PDF search index: discovers every real PDF attachment on disk (via magic-byte detection, not the DB's stored file extension, which can be wrong), extracts text from each, tokenises it, and writes two files:

  • ISE_Data/word_index_pdf.json β€” word β†’ list of [id_attach, page] hits
  • ISE_Data/pdf_store.json β€” id_attach β†’ {filename, pages}

Extraction engines

Runs two independent extractors on every page and takes the union of whatever words each one finds β€” a word either engine catches makes it into the index:

  • PyMuPDF (fitz) β€” fast, handles malformed/corrupted PDFs well
  • pypdf (pinned to 3.1.0 β€” later versions require Python 3.8+ and will crash on this box's Python 3.6)

A third engine, pdfplumber, was tried and rejected β€” its dependency chain requires compiling cryptography, which needs a Rust toolchain not available on this CentOS 6 box. pypdfium2 was also tried and rejected β€” no compatible wheel exists for this platform at any version.

Usage

python3 PDFIndexer.py [options]

Options

FlagDefaultDescription
--attach-dir DIRauto-detected (<board root>/attachments)Directory to scan for PDF attachments
--dry-runoffDiscover + extract only β€” don't write anything to disk
--limit NnoneProcess only the first N discovered PDFs (for testing)
--monitoroffPrint live progress β€” which engine is running per file, new words added per file, and a running summary every 5%
--verboseoffAlso stream log output to stdout, not just the log file
--freshoffIgnore any existing index on disk and rebuild everything from scratch. Default behaviour without this flag is to resume: load the existing index and skip attachments already indexed
--checkpoint-every N25Save progress to disk every N processed PDFs, so a crash (e.g. an OOM kill) only loses recent work rather than the whole run. Set to 0 to disable and only save once at the end
--start-at ID_ATTACHnoneSkip every attachment with id_attach below this value β€” for manually restarting a run from a known point, independent of the resume/skip logic
--fastoffPyMuPDF only β€” skips pypdf entirely. ~20x faster per file and noticeably lighter on memory, at the cost of losing pypdf's independent-parser coverage from the union. Good for quick reindexes or spot-checks

Examples

# Standard run β€” thorough (both engines), resumes automatically if re-run
python3 PDFIndexer.py --monitor

# Quick single-engine pass
python3 PDFIndexer.py --fast --monitor

# Full rebuild from scratch, ignoring whatever's already indexed
python3 PDFIndexer.py --fresh --monitor

# Resume a run that was killed partway, starting from a known attachment ID
python3 PDFIndexer.py --start-at 5235 --monitor

# Test on a handful of files without writing anything
python3 PDFIndexer.py --limit 5 --dry-run --monitor

Notes

  • Memory: this box has 2GB RAM. The full two-engine index is held in memory during a run and has been observed to trigger the Linux OOM killer near the end of a full pass. --checkpoint-every bounds how much work is lost if that happens; --fast reduces peak memory substantially by skipping the second extraction pass. Watch free -h during a full --monitor run if concerned.
  • Tokeniser: currently a plain .split() (tokenise() in the script) β€” not punctuation-aware. [listatt] and listatt are indexed as separate tokens. Flagged as a known limitation, not yet replaced.
  • Filenames: pdf_store.json's filename field is currently always None β€” resolve_filename() exists in the script but isn't wired into main(). PDFsearch.py resolves real filenames itself via a direct DB query instead, so this doesn't block search, but it's a known gap in the index file itself.

PDFsearch.py

Searches the index PDFIndexer.py builds, plus (as of the filename/topic feature) scans every PDF attachment's own filename and its linked forum topic's subject line β€” so a document titled with the right words surfaces even if those words never appear in the extracted PDF text itself.

Usage

python3 PDFsearch.py QUERY [options]

QUERY is a required positional argument β€” the search terms.

Options

FlagDefaultDescription
--limit Nnone (no cap)Cap the number of results returned. Omit for every match β€” useful to pass e.g. --limit 20 for a quick SSH sanity check
--html FILEnoneWrite an HTML results report to this path, in addition to (or instead of) console output
--source {ui,ssh}sshPresent so argparse accepts the literal --source=ui flag the web UI bridge passes. ise_trace.py detects UI vs SSH itself by checking sys.argv, independent of this flag

Examples

# Console search, every match
python3 PDFsearch.py listatt

# Capped, for a quick check over SSH
python3 PDFsearch.py heroin --limit 20

# Full HTML report
python3 PDFsearch.py heroin --html /var/www/html/ISE_Data/results.html

What a result can look like

  • Content match β€” the term was found in the PDF's extracted text. Shows the matching page numbers and a link to open the PDF at that attachment.
  • Title/topic match only β€” the term appears in the attachment's filename or its topic's subject line, but not in the extracted text itself. Still shown (never silently hidden) with a "πŸ”Ž Matched via title/topic" block naming the exact term, field, and value that matched.
  • Both β€” a document can have real page-level content hits and a title/topic match; both are shown together.
  • Orphaned attachment β€” if a matched attachment's id_msg doesn't resolve to a real row in the messages table (the post it was attached to is gone), the result shows "No topic found" instead of an Open PDF link, since SMF's dlattach action requires a valid topic to serve the file at all β€” confirmed there's no working fallback for this case.

Wildcard search

As of this session, query terms containing * or ? are matched via fnmatch (case-sensitive, matching this script's existing case-sensitive exact-term lookup) against every key in word_index_pdf and every filename/topic subject, instead of the normal exact-match/ substring lookup. A term with no wildcard characters keeps the original fast path. Query terms come through tokenise()'s plain text.split() here β€” wildcard characters survive that unchanged since .split() doesn't strip anything, unlike MiscSearch.py's regex tokenizer, which needed its pattern extended to preserve them.

syslog note

UI-triggered searches (via qf_PDFsearch_bridge.php) are syslogged by the PHP bridge before this script even runs. SSH-triggered runs bypass that PHP layer entirely β€” without a fix, they'd never reach syslog at all, only ise_trace.py's separate log. PDFsearch.py now syslogs the query itself (tag ise-pdf-search-ssh, message prefixed [SSH]) when it detects it's running via SSH, so both paths end up logged consistently β€” this same gap existed in search.py and was fixed there too.

Fixed bug (confirmed 2026-09-23, fixed in v7.5.2): zero-match return shape

get_ranked_results() returns a bare [] on a genuine zero-match search (when both content_results and title_hits are empty) β€” but the function's own documented contract, and every real caller (main() here, and MegaSearch.py's PDFsearch.get_ranked_results(...) call), unpacks it as a 2-tuple: results, order_mode = get_ranked_results(...). A genuine zero-match PDF search therefore raises ValueError: not enough values to unpack (expected 2, got 0), not the clean "No matches." path MiscSearch.py/search.py take for the same situation.

Real-world impact, traced end-to-end:

  • Standalone/SSH use β€” caught only by main()'s outer try/except, which prints "ISE ERROR β€” see {log path}" and logs a full traceback, instead of the correct "No matches." line. Confusing if you're hand-verifying a fix over SSH (exactly the kind of check done throughout this project's own troubleshooting workflow) β€” a totally normal zero-match search looks identical in the log to a real crash.
  • Via qf_PDFsearch_bridge.php β€” no visible symptom to the end user: the script exits before --html is ever written, so the bridge's fallback branch (json_decode($output, true) on the "ISE ERROR" text β†’ null β†’ empty() is true) shows the same "No results found for: ..." message a real zero-match would show anyway. The bug is real but currently invisible at the UI layer.
  • Via MegaSearch.py β€” caught by that script's own per-source try/except (see MegaSearch.py section above), so the merged result set is unaffected, but every legitimate "PDF had nothing" case gets logged as "MegaSearch.py: pdf source failed" alongside genuine failures β€” noise that could mask a real problem later.

Fix applied in v7.5.2: the zero-match early-return in get_ranked_results() now returns return [], order_mode instead of a bare return [], restoring the documented 2-tuple contract. Delivered as a corrected file; deployment timing to Live is Andrew's call, per standing policy.


orphan_check.py β€” attachment/post integrity report

A standalone diagnostic script, not part of the search UI or any of the five search sources' dropdowns β€” an admin/maintenance tool meant to run on a schedule (cron) rather than by hand each time.

Finds attachments whose id_msg no longer resolves to a real row in smf209_messages β€” i.e. the post an attachment was uploaded against is gone (deleted, moved, or the upload never got attached to a finished post). Deliberately excludes avatars: SMF stores profile avatars as attachments with id_msg=0 by design, which is normal, not orphaning, and would otherwise swamp every report with roughly 360 harmless entries (confirmed on this install).

Keeps a small state file, ISE_Data/orphan_state.json, so each run reports what's new since the last run instead of re-listing the same known orphans every week.

Usage:

python3 orphan_check.py                  # text report to stdout
python3 orphan_check.py --html out.html   # HTML report
python3 orphan_check.py --quiet-if-empty  # print nothing when no new orphans (good for cron)

Fixed bug (confirmed 2026-09-24, fixed in v7.4.2): broken import

Every invocation of this script β€” CLI, --html, and cron β€” failed immediately with:

ImportError: cannot import name 'BOARD_URL' from 'PDFsearch'

Root cause: PDFsearch.py's own settings-loading was refactored onto ise_settings.py (v4.0, 2026-08-21) two days after orphan_check.py's last edit (2026-08-19). That refactor renamed PDFsearch.py's public base-URL constant from BOARD_URL to BASE_URL (now sourced from ise_settings.get_base_url()); orphan_check.py was never updated for the rename and kept importing the old name.

Fix applied in v7.4.2: BOARD_URL was never actually referenced anywhere else in this file (confirmed β€” it builds no links and does nothing with a base URL), so rather than re-import it under its new name only to leave it unused, it was simply dropped from the import line, leaving from PDFsearch import DB_CONFIG, ISE_ROOT β€” both of which this file does use. Delivered as a corrected file; deployment timing to Live is Andrew's call, per standing policy.


pdf_duplicate_spotter.py

Finds byte-identical PDF duplicates in the attachments directory β€” the same file content re-uploaded under a different id_attach, usually with a different filename and/or topic. Not near-duplicate detection (a different scan of the same book, an edited copy, etc.) β€” that's a separate, harder problem this doesn't attempt.

Reuses existing infrastructure rather than reimplementing it: discover_pdfs()/ATTACH_DIR from PDFIndexer.py for the same magic-byte PDF scan already used for indexing, and DB_CONFIG/get_topic_for_attachment()/build_download_url() from PDFsearch.py for the same DB lookups/link-building already used there.

Method: SHA-256 over each file's actual bytes, streamed in 1MB chunks (2GB-RAM box constraint β€” never loads a whole PDF into memory). Files sharing a hash are byte-identical, full stop β€” no false positives possible from this method, and correspondingly no way to catch a duplicate that isn't byte-identical (see PIHKAL/TIHKAL note below).

Usage

python3 pdf_duplicate_spotter.py [options]

Options

FlagDefaultDescription
--attach-dir DIRATTACH_DIR (from PDFIndexer.py)Directory to scan for PDFs
--json PATHISE_Data/pdf_duplicates.json (via ISE_ROOT, same sibling-of-The_ISE_Project convention as every other script's data files)Write the full report as JSON here β€” this is what pdf_dupe_review.php reads, so leave it at the default unless you're also pointing that page somewhere else

Examples

# Standard run β€” writes to the default path pdf_dupe_review.php expects
python3 pdf_duplicate_spotter.py

# Custom output location
python3 pdf_duplicate_spotter.py --json /tmp/dupes.json

Companion page: pdf_dupe_review.php

Not a Python script (out of this manual's stated scope, noted here for findability) β€” an admin-only page in The_ISE_Project/ that reads this script's --json output and lists every duplicate copy as two plain links: View post (opens the actual forum post) and Delete post (single click, JS-confirmed, SMF session-token guarded, deletes the whole post via SMF's own removeMessage() β€” which also cleans up that post's attachment as part of the same call). Backs up the message row to ise_pdf_dupe_backup_v2 before deleting (metadata only, not the PDF bytes β€” a forensic record, not a one-click undo, since SMF reassigns ids on delete) and remembers deleted ids in ISE_Data/pdf_dupe_deleted.json so they stop appearing without needing to rerun the spotter.

Notes

  • Exact-match only, by design. Andrew's PIHKAL/TIHKAL attachments didn't show up in an early run β€” expected, since those are almost certainly different scans/uploads of the same book rather than byte-identical files. Confirmed as the correct behavior for the scope chosen (exact duplicates), not a bug.
  • Orphaned attachments (no id_topic/id_msg β€” the post they were attached to is already gone) still appear in a duplicate group if their bytes match, but pdf_dupe_review.php has no post to link to or delete for those β€” they show with no View/Delete links.

MiscIndexer.py

Builds the search index for text-based attachments: .md, .txt, .text, .json, and .html (full content, tokenised and indexed). Simpler than PDFIndexer.py in a few key ways β€” these are already plain text, so there's no extraction step, no multi-engine union, and no resume/checkpoint logic (every run is a full fresh rebuild, since the corpus is expected to stay small).

.html gets one extra step before tokenising: strip_html_for_indexing() removes <script>/<style> blocks (tag and content together), strips every remaining tag, and decodes HTML entities (&amp; β†’ &, etc.) via re-based regexes β€” deliberately simple, not a full HTML parser, good enough for indexing purposes only. This keeps markup (div, class, href, inline JS/CSS) out of word_index_misc.json β€” only visible page text is indexed. misc_store.json's content field for .html entries holds the stripped text (so search snippets are readable), but the checksum field is still computed on the original, unstripped file.

MiscIndexer.py also discovers a much wider set of code/config attachment types β€” .py, .php, .js, .sh, .sql, .css, .c, .cpp, .java, .rb, .go, .rs, .pl, .xml, .yml, .yaml, .ini, .conf β€” but treats them as filename/type-only: FILENAME_ONLY_EXTENSIONS skips read_text_file() and tokenise() entirely for these, so nothing they contain enters word_index_misc.json. Their misc_store.json entry has content: "", zero word counts, and filename_only: true β€” they're viewable via the "Open file" download link and filename/topic searchable, but never full-text searched. .html is deliberately not in this set β€” it's mostly prose with markup wrapped around it, so it gets full content indexing (above) instead.

  • ISE_Data/word_index_misc.json β€” word β†’ list of {id_attach, count} hits
  • ISE_Data/misc_store.json β€” id_attach β†’ full metadata + raw content (filename, file_type, board/topic/message context, headings, word count, checksum, filename_only flag, etc. β€” resolved once at index time via a DB join through messages/topics/boards, so MiscSearch.py never needs a separate lookup)

Discovery

Unlike PDFIndexer.py's magic-byte scan, plain-text files have no reliable magic number to sniff β€” and SMF strips real file extensions from attachments' on-disk filenames anyway. Discovery is therefore DB-driven by necessity: MiscIndexer.py cross-references smf_attachments.fileext against what's actually present on disk (matched by numeric id_attach prefix), and only indexes attachments that exist in both places.

Usage

python3 MiscIndexer.py [options]

Options

FlagDefaultDescription
--attach-dir DIRauto-detected (<board root>/attachments)Directory to scan for text attachments
--dry-runoffDiscover + read only β€” don't write anything to disk
--limit NnoneProcess only the first N discovered files (for testing)
--monitoroffPrint live progress and running stats (unique words so far, in-memory KB) every ~5%
--verboseoffAlso stream log output to stdout, not just the log file

Examples

# Standard run β€” full rebuild, verbose progress
python3 MiscIndexer.py --monitor

# Test on a handful of files without writing anything
python3 MiscIndexer.py --limit 5 --dry-run --monitor

End-of-run summary

Always prints (regardless of --monitor): attachments matched vs indexed, counts skipped for non-UTF-8 decoding or being empty, a breakdown by file type, total unique words in the index, average word count per file, on-disk size of the written index files, and the log file path. Non-UTF-8 files are skipped and logged, not decoded with a fallback (no latin-1 attempt) β€” check the log for exact filenames if any show up as skipped.

Notes

  • Every run is a full rebuild. No resume, no incremental mode, no checkpointing β€” deliberate, since the corpus is expected to stay small relative to posts/PDFs. Revisit this (and possibly borrow PDFIndexer.py's resume/checkpoint pattern) if the .md/.txt/ .text/.json attachment count grows substantially.
  • Tokeniser: a simple self-contained regex, not the shared Tokeniser.py module used by post/PDF search β€” a known simplification, not yet reconciled for exact behavioral parity.
  • Headings: extracted from #/##... markdown syntax, only for file_type == "md" β€” always empty for .txt/.text/.json.

MiscSearch.py

Searches the index MiscIndexer.py builds. Also searches filename and topic subject (same rationale as PDFsearch.py's title/topic matching) β€” merged into one result set per attachment, with title-only matches shown alongside content matches rather than silently dropped.

Usage

python3 MiscSearch.py QUERY [options]

Options

FlagDefaultDescription
--limit Nnone (no cap)Cap the number of results returned
--html FILEnoneWrite an HTML results report to this path
--source {ui,ssh}sshSame purpose as PDFsearch.py's flag β€” argparse compatibility for the bridge's --source=ui; ise_trace.py detects UI vs SSH by checking sys.argv directly

Examples

# Console search
python3 MiscSearch.py listatt

# Full HTML report
python3 MiscSearch.py heroin --html /var/www/html/ISE_Data/misc_results.html

Wildcard search

As of this session, query terms containing * (any run of characters) or ? (exactly one character) are matched via fnmatch against every key in word_index_misc (content) and every filename/topic subject (title search), instead of the normal exact-match/substring lookup. A term with no wildcard characters still takes the original fast path β€” wildcards only trigger the full-vocabulary scan when present. The query tokenizer regex (_TOKEN_RE) was extended to preserve */? so they survive tokenisation β€” the content tokenizer in MiscIndexer.py was left unchanged, since indexed file content never legitimately contains literal wildcard characters worth preserving.

Filename-only results

Entries with filename_only: true in misc_store.json (see MiscIndexer.py above) render with a "πŸ“ Filename/type match only" badge in place of a content snippet β€” they can only ever appear via search_filename_topic(), never via content search, since nothing of theirs is in the word index.

Differences from PDFsearch.py

  • No pagination/pages concept β€” a .md/.txt file is one document, not a multi-page PDF, so each result is a single card with a content snippet around the matched term, not a list of matching page numbers.
  • No lazy DB lookup at search time β€” MiscIndexer.py already resolved and stored the owning topic/message/board context at index time, so "Jump to post" links are built straight from misc_store.json with no extra query.
  • Always prints a proper end-of-run summary (query, term list, content vs title match counts, total results, HTML path if applicable) β€” a no-match search shows 0 in that summary rather than a standalone early "No matches." message, matching the pattern MiscIndexer.py already used.

HTML output / theming

Colors are isolated into a single :root { --var: ... } CSS block at the top of the generated HTML, matching the pattern used for PDF search results β€” intended so a theme can be swapped in (e.g. by reskinning the :root block) without touching the card markup/logic underneath it.

syslog note

Same gap and fix as PDFsearch.py/search.py: UI-triggered searches are syslogged by qf_Miscsearch_bridge.php before this script runs; SSH-triggered runs syslog the query themselves (tag ise-misc-search-ssh, message prefixed [SSH]) so both paths are consistently logged.


MegaSearch.py

Queries search.py, PDFsearch.py, and MiscSearch.py's get_ranked_results() directly (no subprocess, no re-searching through their CLI layers) and merges the three result lists into one page. Does not include ISEmedia β€” despite the name, ISEmega is "Posts + PDF + Misc, merged," not "everything." ISEmedia is a fifth, separate, standalone search source with its own dropdown option; nothing currently feeds Media results into a Mega search. Worth knowing plainly since the name invites the assumption otherwise.

Each of the three source calls is individually wrapped in its own try/except Exception, logged via ise_trace.log() as e.g. "MegaSearch.py: pdf source failed" and left as an empty list rather than aborting the whole search β€” so one source being unavailable (a DB connection failure, an index file missing) degrades gracefully instead of taking out the other two. Known caveat, not yet fixed: that same try/except also catches a completely normal zero-match PDF result (see PDFsearch.py's "Known bug" note below) β€” a genuine "PDF just didn't match anything" case gets logged identically to a real failure ("pdf source failed"), which is misleading in the trace log even though the merged Mega result set itself still comes out correct (posts/misc results are unaffected). Fixing the root cause in PDFsearch.py would also clean up this log noise.

Merge strategy

Two different merge strategies depending on order_mode, not one merge with a sort bolted on afterward:

  • relevance/score (default) β€” ise_render.interlace_results(), round-robin across the three tagged lists so one source's higher raw scores can't crowd the other two off the page entirely.
  • newest/oldest β€” ise_render.merge_by_date(), a flat poster_time sort across all three tagged lists. newmod/oldmod are not offered at the Mega level β€” there's no coherent way to merge a posts-only "last edited" concept against two sources that don't have it at all; use ISE (posts) search directly for those.

The /sort/#ise-order- directive is stripped from the query once, in get_mega_results(), before any of the three sources sees it β€” so all three always search on an already-clean query and apply their own default (relevance) sort internally; the redundant per-source sort when a date directive is active is intentional, it guarantees the three sources can never land on inconsistent per-source order_modes.

Result cap β€” the one source that actually has a default cap

Unlike search.py/PDFsearch.py/MiscSearch.py on their own (all effectively unlimited via the UI β€” see each script's own --limit note above; none of their bridges pass --limit), MegaSearch.py is the one source with a real default result cap:

MEGA_SOFT_CAP = 180
MEGA_PAGE_SIZE = None   # argparse default for --limit

run()'s cap precedence: an explicit --limit N always wins; else --all (or &all=1 on the page) bypasses the cap entirely; else MEGA_SOFT_CAP (180) applies. Since qf_MegaSearch_bridge.php passes neither --limit nor --all, every UI-triggered Mega search is capped at 180 displayed results by default β€” everything is still searched and ranked first, the cap only limits how many result cards get drawn, and the on-page notice says so with a "Show all" link.

Usage

python3 MegaSearch.py "your query here" [--html FILE] [--limit N] [--all] [--source {ui,ssh}]

Card rendering β€” delegates rather than duplicates

_render_posts_card() delegates to ResultFormatter.py's real _render_cluster_item() (wrapping each flat post result as a single-item pseudo-cluster) rather than maintaining a hand-copied version β€” this is a deliberate fix for how the Formatted/Raw buttons and ranking explainer drifted out of sync the first time this was tried as a separate implementation. PDF/Misc cards are Mega's own rendering.


MediaIndexer.py / MediaProcessor.py β€” building the ISEmedia index

Two scripts, not one β€” indexing ISEmedia is a two-stage pipeline, unlike the single-script indexers for the other three sources:

  • MediaIndexer.py β€” the cataloger. Scans smf_attachments for image/video extensions (.jpg/.jpeg/.png/.webp/.gif β€” .gif support was added late, 70 real GIFs were silently excluded before that fix), resolves each one's real file path, and writes ISE_Data/media_store.json β€” filename, MIME type (via mimetypes.guess_type()), and forum context (topic/board/post) for every discovered item. Does not compute any embeddings or captions itself β€” that's MediaProcessor.py's job, run second.
  • MediaProcessor.py β€” the heavy-lifting stage. For each item in media_store.json: encodes it with CLIP (ViT-B-32-quickgelu, openai pretrained weights β€” must match MediaSearch.py's CLIP_MODEL_NAME/CLIP_PRETRAINED constants exactly, or query embeddings and stored embeddings come from different vector spaces with no error to indicate it) and, unless --skip-captions is passed, generates a caption via a separate SmolVLM/llama.cpp pipeline (chosen specifically to avoid a glibc/Rust toolchain requirement this box can't meet β€” see the ISEmedia build checkpoints for the full story). For video, samples --samples frames (default 3) from the first --window seconds (default 30.0) rather than processing the whole file. --chunk-size N processes at most N new items and stops (re-run the same command to resume) β€” useful for a long backlog without holding one giant run open. --force/--reindexall re-processes items that already have embeddings, for a model/caption pipeline change.

Both scripts, shared traits

  • --attachid N β€” target one specific attachment, for focused testing.
  • --monitor β€” execution telemetry/performance metrics.
  • --verbose β€” detailed diagnostic logging.
  • Settings: both import get_board_root()/get_data_dir() (and MediaProcessor.py also locate_attachment_file()/ get_attachment_directories()) from Collabware.core_utils, not ise_settings.py β€” see this manual's opening note. mysql_available is probed via a soft try: import pymysql in MediaIndexer.py, matching this project's general pattern of degrading gracefully rather than hard-requiring every optional dependency.
  • Interpreter: both need python3.8 specifically β€” CLIP/torch/ open_clip are only installed under python3.8's site-packages on Live, not the system default python3/python3.6 the rest of ISE's scripts run under (see MediaSearch.py/qf_Mediasearch_bridge.php below for the full story of how that was diagnosed and fixed for the search side; the same interpreter requirement applies here, run these two manually as python3.8 MediaIndexer.py ... / python3.8 MediaProcessor.py ..., not bare python3).

Versioning β€” a different convention from the rest of ISE

Both scripts' docblocks carry a version tag that looks like an overall project version ("ISE v9.7" for MediaIndexer.py, "ISE v9.18" for MediaProcessor.py" as of this writing) plus a RUNTIME_VERSION string kept manually in sync with it β€” this is a **separate, Collabware-native convention**, not the same thing as every other ISE module's own per-file semver (search.py v7.5.2, MediaSearch.py v1.2.0, etc.). Don't read "ISE v9.7" here as meaning the same thing as the package's own "Version 9.5" (install_ise.php`) β€” they're two different counters from two different codebases that happen to share a project.


MediaSearch.py β€” ISEmedia search

v1.7.1 (2026-09-28, DM2S/NX v9.8). Keyword search over images, video and audio. No model is loaded at search time. The AI work (CLIP embeddings, SmolVLM captions, CLAP audio tags) is done at index time by MediaProcessor.py / AudioProcessor.py. CLIP re-ordering stays available as an opt-in (--clip).

What it searches

Candidates are every attachment in media_store.json (MediaIndexer's catalogue, including .mp3/.wav), plus anything in captions_store.json and audio_store.json. A search term is checked against five fields:

FieldSourceMatch rule
FilenameDB a.filename (else media_store)Whole string. A plain word is a substring; a wildcard must fit the whole name
Internal filename<id_attach>_<file_hash> from media_storeSame as filename. Omitted when file_hash is empty
Post subjectDB m.subject, html-unescapedSame as filename
Captioncaptions_store (image, each video frame <id>_t<sec>), audio_store item + windowsWhole words. The best-matching frame or window is shown
Tagsaudio_store item tags ([[keyword, prob, group], ...], keyword used)Whole words

The matching rules are the same as MiscSearch.py's:

  • tokenise() keeps quoted phrases whole and preserves */?.
  • A wildcard is matched with fnmatch on lowercased text.
  • Any term hitting any field includes the item (OR).
  • A quoted phrase is case-insensitive here. Captions have no meaningful case, and Misc's phrase matching is case-sensitive.

Score = the sum of field weights per distinct term hit: caption 40, tags 40, filename 25, internal 25, subject 15. Each result carries match_evidence (term β†’ field β†’ value), which is shown on every card. Avatars are dropped (v1.4.0 rule).

Internal filename (v1.7.1)

The internal name <id_attach>_<file_hash> uses file_hash from media_store.json when it's there. Otherwise it uses a.file_hash from the same DB query that fetches topics. An item known only to audio_store.json or captions_store.json is still found by 6998_*.

Tempo and key filters (v1.7.0)

bpm:124 (Β±1.5), bpm:122-126, and key:Am / key:A-minor / key:F# / key:Gb-major are pulled out of the query before tokenising (parse_filters()).

  • Matching: an item fits a filter when its own value fits, or, for long mixes, when any window does (apply_filters()). The evidence names the windows.
  • Combining: filters are ANDed with each other and with the keyword gate. Words still OR among themselves.
  • Only filters: a query of just filters lists every item that fits.
  • Bad filters: an unusable filter is reported and ignored. A query of only unusable filters returns nothing; it doesn't turn into "list everything".

Features come from audio_store.json (AudioProcessor β‰₯ 1.1.0): bpm, bpm_range, key, lufs, lra and changes per item, plus bpm and key per window. Cards show 118-128 BPM (124), A minor, -9.8 LUFS, ~14 tracks and the approximate track changes.

Usage

python3.8 MediaSearch.py QUERY [--limit N] [--all] [--html FILE] [--clip] [--source {ui,ssh}]
python3.8 MediaSearch.py '*.png'
python3.8 MediaSearch.py '6204_*'
python3.8 MediaSearch.py '"brown horse"'
python3.8 MediaSearch.py '/sort newest *.mp3'

QUERY is taken from argv[0] before argparse sees it, the same as in the other search scripts.

FlagDefaultDescription
--limit NnoneShow exactly the best N
--alloffDraw every match (bypass the soft cap). The bridge passes it for &all=1
--html FILEnoneWrite the results page. Draws at most MEDIA_SOFT_CAP = 180 cards unless --all/--limit is given. When it caps, a red "Showing 180 of N β€” Show all" notice appears (same as ISEmega)
--clipoffLoad CLIP and add similarityΓ—150 to image/video matches. Doesn't change which items match. Reads embeddings_store.json, which nothing else here reads
--keyword-onlyβ€”No-op, kept for compatibility
--source {ui,ssh}sshSame as the other search scripts

The CLI output ends with a Summary: line. An optional store that exists but can't be read is reported with [!] on stderr and as a red line on the results page, and the search carries on without it.

API (for MegaSearch.py)

get_ranked_results(query, limit=None, keyword_only=True) returns (results, order_mode, clip_note, warnings). It was a 3-tuple before v1.6.0. Importing the module no longer imports torch.

Bridge

qf_Mediasearch_bridge.php v1.1.0 no longer passes --limit 180, which silently truncated results. It passes --all when &all=1 is in the URL. It still resolves python3.8 and sets HF_HOME; that's only needed for --clip.

Ranking engine

RankingEngine.py has no media_ranking_engine() (confirmed again against the v9.8 source), so v1.6.0 removed the dead import. Media scoring is the field-weight sum above, which is simple and inspectable.

syslog / help / theming

These are unchanged: /help serves the shared ISE_help.html (its ISEmedia section is updated for v1.6.0), ise_trace logging runs throughout, and the theme key is ise_media_theme.

Query parsing internals (search.py / QueryParser.py / Tokeniser.py / ANDMatcher.py / IndexLookup.py)

For the user-facing summary of what each quoting style does per source, see the shared help page rather than duplicating it here: <boardurl>/The_ISE_Project/ISE_help.html (or type /help in any search box). What follows is the code-level "why" β€” the actual functions and matching paths involved, for anyone debugging or extending this.

search.py (ISE posts) is the only one of the three search sources with a real quoting/phrase concept β€” PDFsearch.py and MiscSearch.py have no equivalent at all (see each script's Wildcard search note above for how they actually handle a literal quote character passed to them: PDFsearch.py's .split() glues quotes onto the token, breaking the match; MiscSearch.py's tokenizer regex silently drops quote characters entirely).

For ISE (posts), the pipeline is QueryParser.parse() β†’ Tokeniser.tokenise() β†’ ANDMatcher.match() β†’ IndexLookup.py's lookup_word()/lookup_phrase():

  • No quotes β€” a normal WORD token. Lowercased, looked up directly against word_index.json (fast path) if it's a plain \w+ string.
  • "double quotes" β€” parsed as a real PHRASE token. Matched via lookup_phrase(): an exact, case-sensitive, literal substring search against the raw post text (phrase in post["text"]) β€” whitespace and punctuation inside the quotes must match exactly.
  • 'single quotes' β€” not special-cased anywhere in this pipeline. QueryParser.py's regex only recognizes "..." for phrases; a single-quoted word is caught by the generic \S+ word pattern including its quote characters. Tokeniser.py's docstring claims WORD terms get "punctuation stripped," but the actual code doesn't strip anything β€” it only lowercases. So 'test' becomes the literal token 'test', which isn't \w+, so IndexLookup.lookup_word() falls through to its punctuation-aware path: a case-insensitive literal substring search for 'test', apostrophes included. In practice this means single-quoting a word in ISE search will almost always return nothing, since real post text essentially never contains that exact quoted form. This is a real, confirmed gap, not a design choice β€” worth fixing (either treat single quotes like double quotes, or strip them) if/when the /help page's quote-behavior explanation needs to describe intended rather than actual behavior.
  • A word containing other punctuation (e.g. claude.ai) β€” same lookup_word() fallback as above: case-insensitive literal substring match against post text, requiring that exact adjacent string β€” not claude and ai matched as separate words elsewhere in the post.

SortEngine.py (shared /sort directive parsing)

Built in v9.3.0. A small shared module β€” extract_order_mode(query, valid_modes, default) β€” used by search.py, PDFsearch.py, MiscSearch.py, and MegaSearch.py so the /sort <mode> and #ise-order-<mode> query-string syntax is recognised identically everywhere, instead of each module hand-rolling its own copy of the same regex (the old version of this file predated RankingEngine.py/ ISEmega entirely and was dead code β€” nothing in the current pipeline called it).

def extract_order_mode(query: str, valid_modes: set, default: str) -> tuple:
   # strips a /sort <mode> or #ise-order-<mode> directive out of query,
   # resolving it against the caller's own valid_modes set.
   # Returns (cleaned_query, order_mode).

/sort is checked first (the documented, user-facing syntax); #ise-order- is the older internal syntax, still recognised for back-compat. If both are present in one query, /sort wins. A mode name not in the caller's valid_modes is left in the query untouched β€” it just fails to match anything as a search term, same as any other typo, rather than silently vanishing with no explanation.

Where it's called, and with what valid_modes:

Callervalid_modesNotes
search.py (ISE/posts)relevance, score, date, newest, oldest, newmod, oldmod, subject, board, author, occurrencesFull ORDER_KEYS dict; date is a legacy alias for newest. Called from inside get_ranked_results() itself (not just run()) so MegaSearch.py's direct call gets directive parsing too β€” before this, a /sort directive routed through ISEmega leaked into the query as literal search terms.
PDFsearch.pyrelevance, score, newest, oldestNo newmod/oldmod β€” see below.
MiscSearch.pyrelevance, score, newest, oldestSame.
MegaSearch.pyrelevance, score, newest, oldestStripped once in get_mega_results(), before any of the three sources is queried, so all three see an already-clean query and never re-derive the directive independently. Does not cover ISEmedia β€” see MegaSearch.py section above.
MediaSearch.pyrelevance, score, newest, oldestSame shape as PDF/Misc β€” no newmod/oldmod (attachments carry poster_time only). The real difference for Media isn't in this parsing step at all β€” it's that a non-empty query also gates results by literal keyword match before any ORDER_KEYS sort runs (v1.2.0), which none of the other four sources need to do here since their own word-index lookup already did that filtering upstream. Note: SortEngine.py's own module docblock (v8.0, 2026-09-14) only names search.py, PDFsearch.py, and MiscSearch.py as callers in its "Used by" line β€” stale as of MediaSearch.py's build (2026-09-23); the real import (from SortEngine import extract_order_mode) confirms Media uses it too, the docblock just predates that.

Why newmod/oldmod is ISE (posts)-only: PDFsearch.py/ MiscSearch.py results carry poster_time only β€” PDF/Misc attachments aren't edited the way a forum post is, so there's no modified_time field in pdf_store.json/misc_store.json to sort by. search.py's ORDER_KEYS["newmod"]/["oldmod"] fall back to poster_time for any post that's never been edited (_effective_modified()), matching ISE_help.html's documented behavior exactly.

ISEmega's date sort is a different merge, not a re-sort: ise_render.py gained merge_by_date() alongside its existing interlace_results(). MegaSearch.get_mega_results() picks one or the other based on order_mode β€” interlace_results() (round-robin, unchanged) for relevance/score, merge_by_date() (flat poster_time sort across all three tagged lists) for newest/oldest. These are genuinely different strategies, not the same merge with a final re-sort bolted on: interlacing exists specifically to stop one source dominating a relevance-score merge, which isn't a concern once you're sorting by date β€” so a date sort bypasses interlacing entirely.

In-app help (/help and the footer link)

ISE_help.html is a single static file living in The_ISE_Project/ (same directory as all five search scripts, web-accessible via the .htaccess <Files "*.html"> rule) β€” the single shared source all five search UIs point to. It now also documents ISEmega and ISEmedia (added this audit β€” both were previously undocumented there despite ISEmega having existed for some time and being partially wired into the page's /sort table already).

search.py, PDFsearch.py, and MiscSearch.py each define the same pair of names near its top (MegaSearch.py reuses search.py's copy directly rather than redefining it β€” see its own /help short-circuit; MediaSearch.py defines its own, same shape):

HELP_TRIGGER = "/help"
HELP_HTML_PATH = os.path.join(os.path.dirname(os.path.abspath(__file__)), "ISE_help.html")

def _serve_help(html_path):
   with open(HELP_HTML_PATH, "r", encoding="utf-8") as f:
       content = f.read()
   if html_path:
       with open(html_path, "w", encoding="utf-8") as f:
           f.write(content)
   return True

At the very top of each script's argument handling, before any real search logic runs, the query is checked against HELP_TRIGGER case/whitespace-insensitively (query.strip().lower() == HELP_TRIGGER). On a match, _serve_help() copies ISE_help.html's content directly into whatever --html path the PHP bridge passed β€” since every bridge already just readfile()s that path (see qf_search_bridge.php/ qf_PDFsearch_bridge.php/qf_Miscsearch_bridge.php), no PHP changes were needed at all. The console/SSH path (no --html) just prints the shared file's own path instead of writing anywhere. Both paths log via ise_trace.log() and, for SSH runs, call ise_trace.new_trace() same as a normal search would.

ISE_help.html carries its own copy of ResultFormatter.py's 13-theme CSS block and switcher, and on load reads whichever of ise_search_theme/ise_pdf_theme/ise_misc_theme is set in localStorage, so it opens already matching whichever theme was last picked on any results page.

Separately, ResultFormatter.py's format_html(), PDFsearch.py's write_html(), and MiscSearch.py's write_html() each got a small permanent footer link β€” <a href="{BASE_URL}/The_ISE_Project/ISE_help.html"> β€” added just before </body>, so the help page is reachable from every normal results page too, not only via the /help keyword.

Planned work (not yet built)

  • Wildcard (*/?) support for ISE (posts) β€” PDFsearch.py and MiscSearch.py have it (see their Wildcard search sections above); search.py doesn't yet, since its matching runs through the modules documented above rather than a direct dict lookup.

Shared dependency notes

No longer one shared runtime β€” two, as of ISEmedia. Everything below the line was true, and mostly still is, for the original four posts/PDF/misc/mega scripts. ISEmedia's two build-side scripts (MediaIndexer.py, MediaProcessor.py) and its search script (MediaSearch.py) need Python 3.8 specifically β€” torch/open_clip are only installed under python3.8's site-packages on Live, and will fail to import under the system default. Confirmed live (2026-09-23): running any of the three Media scripts under bare python3 fails on import numpy/import torch before ever reaching real logic. qf_Mediasearch_bridge.php resolves python3.8 explicitly for this reason (see MediaSearch.py section above); install_ise.php v9.5 does the same for MediaIndexer.py's background install-time job. IndexBuilder.py/PDFIndexer.py/MiscIndexer.py are deliberately not touched to require 3.8 β€” no evidence they need it, and changing their interpreter risks three scripts that already work.

  • Python (search.py / PDFsearch.py / MiscSearch.py / MegaSearch.py / IndexBuilder.py / PDFIndexer.py / MiscIndexer.py): 3.6, CentOS 6. Several package versions are pinned specifically for this β€” do not let pip install --upgrade move past these without re-testing:
    • pymupdf==1.19.2 (later versions have no cp36 wheel)
    • pypdf==3.1.0 (5.x uses typing.Protocol, unavailable before Python 3.8, and will crash immediately on import)
    • mysql-connector-python β€” used by MiscIndexer.py/MiscSearch.py for the DB-driven discovery/context-resolution join
  • Python (MediaIndexer.py / MediaProcessor.py / MediaSearch.py): 3.8 specifically, on the same CentOS 6 box β€” torch, open_clip, and numpy are the load-bearing reason; see above.
  • DB access / settings: search.py/PDFsearch.py/MiscSearch.py/ MegaSearch.py/MediaSearch.py resolve their base URL via ise_settings.py's Settings_ISE.json lookup (see this manual's opening note β€” this replaces an earlier "via ISE_ROOT" description that predates ise_settings.py v4.0). MediaIndexer.py/ MediaProcessor.py resolve settings via Collabware.core_utils instead β€” a separate mechanism, see the MediaIndexer.py/ MediaProcessor.py section above. DB credentials themselves still come from Settings.php for all scripts; no separate config file or hardcoded credentials anywhere.
  • ise_trace.py: shared logging/tracing module imported by all five search scripts (search.py, PDFsearch.py, MiscSearch.py, MegaSearch.py, MediaSearch.py) and the post/PDF/Misc indexers. MediaIndexer.py/MediaProcessor.py do not import it β€” they log through their own Collabware-native telemetry (--monitor/--verbose) instead, consistent with those two scripts sitting outside the ise_trace-sharing group. On an SSH-triggered run, new_trace() starts a fresh block in the shared ISE_Data/ise_trace.log; a UI-triggered run (detected via the bridge's --source=ui) continues the existing request's trace instead. MiscIndexer.py always starts a fresh trace block regardless, since it's manual/SSH-only β€” no bridge ever invokes it. All five search scripts wrap main()/run() in a top-level exception handler that logs the full traceback via ise_trace.log_exception() before exiting non-zero.
  • syslog vs ise_trace.py β€” not the same log. ise_trace.py's log is detailed pipeline tracing for debugging; syslog (facility LOCAL0) is the actual search-query audit trail reviewed via daily_queries.py/ query_log_reader.py. Historically only UI-triggered searches reached syslog (via the PHP bridge/Search.php) β€” SSH-triggered runs of search.py/PDFsearch.py/MiscSearch.py silently never syslogged at all until this was caught and fixed across all three. MediaSearch.py logs UI-triggered queries via qf_Mediasearch_bridge.php's own syslog() call, same pattern.


Claude: [28/09/2026--06:34:21:321 AEST]

Here's all the QAT work in one bundle: DM2S-NX_QAT_softcopy_v1.1.0.zip, 19 files.

FolderWhat's in it
tools/ise_qat.py v1.1.0 (the battery) and ise_forge.py v1.0.0 (backward tests), ready for /bin
docs/both manuals, plus my_queries.txt.sample
examples/a real QAT report, a forge score sheet, and the AI prompt, all from the test run
harness/the stand-in forum data and fake database I tested with, for reference
README.mdinstall steps and what the battery found

The two tools are byte-for-byte the same as the ise_qat_v1.1.0.zip I sent yesterday morning. That one was just easy to miss among everything else. Nothing has changed since; this bundle only adds the examples, the harness and a README.




How to download youtube movies as MP3 or MP4:

πŸ“ Inline Markdownv4.5 · 2026-09-28

PS C:\Users\Admin>Β  Β  yt-dlp --version 2026.08.19

PS C:\Users\Admin>Β  Β  deno --version deno 2.9.7 (stable, release, x86_64-pc-windows-msvc) v8 15.0.245.2-rusty typescript 6.0.3

PS C:\Users\Admin>Β  Β  ffmpeg -version ffmpeg version N-125875-g5d4d3bdc61-20260731 Copyright (c) 2000-2026 the FFmpeg developers built with gcc 15.2.0 (crosstool-NG 1.28.0.23_185f348) configuration: --prefix=/ffbuild/prefix --pkg-config-flags=--static --pkg-config=pkg-config --cross-prefix=x86_64-w64-mingw32- --arch=x86_64 --target-os=mingw32 --enable-gpl --enable-version3 --disable-debug --disable-w32threads --enable-pthreads --enable-iconv --enable-zlib --enable-libxml2 -- enable-libvmaf --enable-fontconfig --enable-libharfbuzz --enable-libfreetype --enable-libfribidi --enable-vulkan --enable-libvorbis --disable-libxcb --disable-xlib --disable-libpulse --enable-gmp --enable-lzma --enable-liblcevc-dec --enable-opencl --enable-amf --enable-libaom --enable-libaribb24 --enable-avisynth --enable-chromaprint --enable-libdav1d --enable-libdavs2 --enable-libdvdread --enable-libdvdnav --disable-libfdk-aac --enable-ffnvcodec --enable-cuda-llvm --enable-frei0r --enable-libgme --enable-libkvazaar --enable-libaribcaption --enable-libass --enable-libbluray --enable-libjxl --enable-libmp3lame --enable-libopus --enable-libplacebo --enable-librist --enable-libssh --enable-libtheora --enable-libvpx --enable-libwebp --enable-libzmq --enable-lv2 --enable-libvpl --enable-openal --enable-liboapv --enable-libopencore-amrnb --enable-libopencore-amrwb --enable-libopenh264 --enable-libopenjpeg --enable-libopenmpt --enable-librav1e --enable-librubberband --enable-schannel --enable-sdl2 --enable-libsnappy --enable-libsoxr --enable-libsrt --enable-libsvtav1 --enable-libtwolame --enable-libuavs3d --disable-libdrm --enable-vaapi --enable-libvidstab --enable-libvvenc --disable-whisper --enable-libx264 --enable-libx265 --enable-libxavs2 --enable-libxvid --enable-libzimg --enable-libzvbi --extra-cflags=-DLIBTWOLAME_STATIC --extra-cxxflags= --extra-libs=-lgomp --extra-ldflags=-pthread --extra-ldexeflags= --cc=x86_64-w64-mingw32-gcc --cxx=x86_64-w64-mingw32-g++ --ar=x86_64-w64-mingw32-gcc-ar --ranlib=x86_64-w64-mingw32-gcc-ranlib --nm=x86_64-w64-mingw32-gcc-nm --extra-version=20260731 libavutilΒ  Β  Β  61.Β  5.100 / 61.Β  5.100 libavcodecΒ  Β   63.Β  7.100 / 63.Β  7.100 libavformatΒ  Β  63.Β  5.101 / 63.Β  5.101 libavdeviceΒ  Β  63.Β  2.100 / 63.Β  2.100 libavfilterΒ  Β  12.Β  3.101 / 12.Β  3.101 libswscaleΒ  Β   10.Β  2.100 / 10.Β  2.100 libswresampleΒ   7.Β  2.100 /Β  7.Β  2.100  Exiting with exit code 0

PS C:\Users\Admin> mkdir C:\yt

Β  Β  Directory: C:\

ModeΒ  Β  Β  Β  Β  Β  Β  Β   LastWriteTimeΒ  Β  Β  Β   Length Name ----Β  Β  Β  Β  Β  Β  Β  Β   -------------Β  Β  Β  Β   ------ ---- d-----Β  Β  Β  Β  26/09/2026Β   6:51 PMΒ  Β  Β  Β  Β  Β  Β  Β  yt

#

DL AUDIO here:

PS C:\Users\Admin> cd C:\yt PS C:\yt> yt-dlp -x --audio-format mp3 --audio-quality 0 "[youtube]https://youtu.be/5YewUmayUrs[/youtube]"

[youtube] Extracting URL: [youtube]https://youtu.be/5YewUmayUrs[/youtube] [youtube] 5YewUmayUrs: Downloading webpage [youtube] 5YewUmayUrs: Downloading visionos player API JSON [youtube] 5YewUmayUrs: Downloading m3u8 information [info] 5YewUmayUrs: Downloading 1 format(s): 251 [download] Destination: Kylie Minogue - Enjoy Yourself (Official Audio) [5YewUmayUrs].webm [download] 100% ofΒ  Β  3.39MiB in 00:00:03 at 940.14KiB/s [ExtractAudio] Destination: Kylie Minogue - Enjoy Yourself (Official Audio) [5YewUmayUrs].mp3 Deleting original file Kylie Minogue - Enjoy Yourself (Official Audio) [5YewUmayUrs].webm (pass -k to keep)

#

DL VIDEO here:

PS C:\yt> yt-dlp -x --audio-format mp3 --audio-quality 0 "[youtube]https://youtu.be/M1fnv9ULgLg?s[/youtube]i=ajZa4Hs31UJYCDXE" [youtube] Extracting URL: [youtube]https://youtu.be/M1fnv9ULgLg?s[/youtube]i=ajZa4Hs31UJYCDXE [youtube] M1fnv9ULgLg: Downloading webpage [youtube] M1fnv9ULgLg: Downloading visionos player API JSON [youtube] M1fnv9ULgLg: Downloading m3u8 information [info] M1fnv9ULgLg: Downloading 1 format(s): 251 [download] Destination: Bacchanaliaβ€˜s - Armageddon - β€˜The Ultimate Encounter’ @ the Hordern Pavilion May 26, 1990 [M1fnv9ULgLg].webm [download] 100% ofΒ  Β  8.44MiB in 00:00:08 at 1.04MiB/s [ExtractAudio] Destination: Bacchanaliaβ€˜s - Armageddon - β€˜The Ultimate Encounter’ @ the Hordern Pavilion May 26, 1990 [M1fnv9ULgLg].mp3 Deleting original file Bacchanaliaβ€˜s - Armageddon - β€˜The Ultimate Encounter’ @ the Hordern Pavilion May 26, 1990 [M1fnv9ULgLg].webm (pass -k to keep) PS C:\yt>Β  Β  yt-dlp -S "vcodec:h264,res,acodec:m4a" --merge-output-format mp4 "[youtube]https://youtu.be/5YewUmayUrs[/youtube]" [youtube] Extracting URL: [youtube]https://youtu.be/5YewUmayUrs[/youtube] [youtube] 5YewUmayUrs: Downloading webpage [youtube] 5YewUmayUrs: Downloading visionos player API JSON [youtube] 5YewUmayUrs: Downloading m3u8 information [info] 5YewUmayUrs: Downloading 1 format(s): 137+140 [download] Destination: Kylie Minogue - Enjoy Yourself (Official Audio) [5YewUmayUrs].f137.mp4 [download] 100% ofΒ  Β  5.54MiB in 00:00:09 at 580.68KiB/s [download] Destination: Kylie Minogue - Enjoy Yourself (Official Audio) [5YewUmayUrs].f140.m4a [download] 100% ofΒ  Β  3.47MiB in 00:00:05 at 634.96KiB/s [Merger] Merging formats into "Kylie Minogue - Enjoy Yourself (Official Audio) [5YewUmayUrs].mp4" Deleting original file Kylie Minogue - Enjoy Yourself (Official Audio) [5YewUmayUrs].f137.mp4 (pass -k to keep) Deleting original file Kylie Minogue - Enjoy Yourself (Official Audio) [5YewUmayUrs].f140.m4a (pass -k to keep)

#

DL VIDEO here:

PS C:\yt> PS C:\yt>Β  Β  yt-dlp -S "vcodec:h264,res,acodec:m4a" --merge-output-format mp4 "[youtube]https://youtu.be/M1fnv9ULgLg?s[/youtube]i=ajZa4Hs31UJYCDXE"

[youtube] Extracting URL: [youtube]https://youtu.be/M1fnv9ULgLg?s[/youtube]i=ajZa4Hs31UJYCDXE [youtube] M1fnv9ULgLg: Downloading webpage [youtube] M1fnv9ULgLg: Downloading visionos player API JSON [youtube] M1fnv9ULgLg: Downloading m3u8 information [info] M1fnv9ULgLg: Downloading 1 format(s): 135+140 [download] Destination: Bacchanaliaβ€˜s - Armageddon - β€˜The Ultimate Encounter’ @ the Hordern Pavilion May 26, 1990 [M1fnv9ULgLg].f135.mp4 [download] 100% ofΒ   51.34MiB in 00:01:34 at 557.32KiB/s [download] Destination: Bacchanaliaβ€˜s - Armageddon - β€˜The Ultimate Encounter’ @ the Hordern Pavilion May 26, 1990 [M1fnv9ULgLg].f140.m4a [download] 100% ofΒ   10.43MiB in 00:00:28 at 368.64KiB/s [Merger] Merging formats into "Bacchanaliaβ€˜s - Armageddon - β€˜The Ultimate Encounter’ @ the Hordern Pavilion May 26, 1990 [M1fnv9ULgLg].mp4" Deleting original file Bacchanaliaβ€˜s - Armageddon - β€˜The Ultimate Encounter’ @ the Hordern Pavilion May 26, 1990 [M1fnv9ULgLg].f140.m4a (pass -k to keep) Deleting original file Bacchanaliaβ€˜s - Armageddon - β€˜The Ultimate Encounter’ @ the Hordern Pavilion May 26, 1990 [M1fnv9ULgLg].f135.mp4 (pass -k to keep)

PS C:\yt>

« Last Edit: Today at 06:15:28 AM by Chip »
friendly
0
funny
0
informative
0
agree
0
disagree
0
like
0
dislike
0
No reactions
No reactions
No reactions
No reactions
No reactions
No reactions
No reactions
Our Discord Server invitation link is https://discord.gg/jB2qmRrxyD

Tags:
 


dopetalk does not endorse any advertised product nor does it accept any liability for it's use or misuse





TERMS AND CONDITIONS

In no event will d&u or any person involved in creating, producing, or distributing site information be liable for any direct, indirect, incidental, punitive, special or consequential damages arising out of the use of or inability to use d&u. You agree to indemnify and hold harmless d&u, its domain founders, sponsors, maintainers, server administrators, volunteers and contributors from and against all liability, claims, damages, costs and expenses, including legal fees, that arise directly or indirectly from the use of any part of the d&u site.


TO USE THIS WEBSITE YOU MUST AGREE TO THE TERMS AND CONDITIONS ABOVE


Founded December 2014
SimplePortal 2.3.6 © 2008-2014, SimplePortal