Skip to content

Repository files navigation

RDP Web

RDP Web is a browser-native port of the Recombination Detection Program workflow. It is designed as a static site: alignments are parsed and analysed locally in a Web Worker, while the numerical core runs in WebAssembly.

Session 26 source checkpoint — RDP-only cyclic parity audit. This source-only archive retains Session 25's fragment evidence and Pages correction, ports CentreBP coordinate centering, BestXOList-style committed-event caching, the AddjustCXO/MakePairsP DoPairs rescan gate, final-RList erasure, exact active role-score rounding, and a graceful 64-round diagnostic ceiling. Full desktop cyclic parity remains unclaimed.

The current native-parity preview additionally reproduces the installed 5.93 zero-bootstrap event-tree call, its exact Clearcut/Tree2ArrayP2 matrices, and an expanded but still incomplete MakeConsensusC statistic set. Dataset0 currently yields 53 Web events versus 48 desktop events; this deployed checkpoint is intended for comparative testing, not parity-sensitive production use.

What this checkpoint contains

  • The RDP5 dataset → settings → primary scan → event reconciliation → ordered review → export workflow.
  • A purely cosmetic Windows 95 application skin: teal desktop, navy title bar, gray beveled controls, recessed data panes, system-font typography, menu/status chrome, and segmented scan progress. It adds no external font/image dependency and changes no analytical computation.
  • FASTA, GDE, CLUSTAL/MUSCLE, sequential/interleaved PHYLIP, NEXUS, and MEGA alignment readers.
  • Alignment diagnostics, pairwise identity checks, and the supplied RDP5 auto-mask workflow.
  • The manual's distinct enabled/masked/disabled row states: masked sequences skip primary triplets but remain eligible for secondary evidence and tree placement; disabled sequences skip primary and secondary evidence while remaining available as phylogenetic context. Modern bulk controls restore the supplied auto-mask, enable all, mask all, or disable all without rendering every row.
  • The manual's automated query-vs-reference workflow. A dataset-level role editor accepts explicit numeric reference groups or detects documented REF-A<name>-style prefixes. Filter-aware multi-row selection can assign or clear a group without rendering every selected row, and group IDs can be compacted by first input appearance without changing membership. The primary scheduler then lazily emits exactly one query plus two references from different groups. Reference records remain eligible to be inferred as recombinants, as the manual requires. The exact record-triplet workload and the supplied reference-group pairs × queries correction factor remain separate and visible, avoiding an O(T) materialized analysis list on large inputs. Reference-as-recombinant calls receive the manual's distinct amber review treatment and explicit JSON/CSV input-role data.
  • A C++20/WASM primary RDP scanner derived only from the supplied RDP sources: information-rich triplet sites, rolling pair counts, candidate tract boundaries, the binomial tail calculation, 169-site scaling, RDP5's multiple-testing cap, and post-score CentreBP midpoint coordinates with source-shaped post-erasure missing-site relocation.
  • A second, independently auditable C++20/WASM discovery stream for the supplied MaxChi MCXoverF workflow: all three variable-site pair profiles, native critical-difference screening, raw strongest-peak order, source-shaped window growth, FindSide, OptLeftBPMC/ OptRightBPMC, tract construction, literal twelve-term/eleven-divisor SmoothChiValsP, completed/rejected peak destruction, and the supplied three-wasted/100-attempt retry bounds. A linear-time heap build plus lazy destroyed-peak rejection preserves the source's chi/boundary/pair order without repeatedly rescanning every surviving profile cell.
  • A third independently labelled discovery stream for the supplied CHIMAERA workflow. Every triplet member rotates through the candidate-recombinant role; monomorphic/all-different sites are removed; the target is encoded as a binary match to either parent; and raw χ² peaks enter the same source-shaped growth, tract-side, boundary-optimization, destruction, and retry lifecycle. The supplied default is 60 information-rich sites. RDP categories, MaxChi equality tracks, CHIMAERA target inputs, and the triplet missing/erasure map are prepared in one alignment-byte pass.
  • A fourth independently labelled discovery stream for the supplied ordinary automated GENECONV workflow. The shared non-monomorphic profile becomes three inner pair-match and three outer discordant-sequence tracks; source mismatch penalties, lambda/K calculation, strict critical scores, Karlin–Altschul tails, six-track provisional roles, stable lowest-P selection, and the configured overlap rule are retained. The active automated defaults are ignored indels, G=1, one overlap, and inactive minimum-fragment controls. Prefix/next-lower/rightmost-maximum queries preserve the start-extension stop/tie rules in O(R log R) per track, and lazy range coverage keeps fragment selection bounded without another alignment scan.
  • A fifth independently labelled discovery stream for the supplied automated 3SEQ workflow. Each triplet member rotates through the candidate-recombinant role in source order; only sites where the parents differ and the target matches exactly one are retained; and the source maximum descent/ascent plus CheckwrapC tract construction is preserved. A compact exact hypergeometric random-walk dynamic program replaces the desktop four-dimensional lookup table within a bounded transition budget, with the supplied SiegmundDiscrete and scaled-exact routes explicitly labelled beyond it. The pre-CheckwrapC probability excursion is retained separately from any larger origin-extended boundary excursion, matching the supplied call order. 3SEQ applies the source's p > 10^-15 Dunn–Šidák / smaller-tail product branch and adds no alignment-byte pass because it reuses the MaxChi/CHIMAERA equality profile.
  • A sixth independently labelled discovery stream for the supplied automated distance-mode BootScan workflow. It preserves seeded SEQBOOT2 weights, replicate-zero unresampled windows, FastBootDist Jukes–Cantor saturation, strict GetPltVal closest-pair votes, source-shaped circular support-region discovery, and separate bootstrap-support, raw MakeScoresBS binomial, and project-corrected probabilities. Its bounded 64 MiB pair-profile cache reuses shared pair/window/bootstrap distances across triplets and is invalidated after every cyclic erasure.
  • A seventh independently labelled discovery stream for the supplied SISCAN workflow. It preserves nearest-fourth-sequence GetSSOL selection on a round-cached source WPGMA/cophenetic context, gap-stripped one/two/three-variable pattern categories, the supplied QuickCheckB fast-screen control-flow quirk, Microsoft-CRT MakeVRand flat-prefix vertical permutations, pair-switch region construction, ShrinkRegionC bounds, and distinct normal-tail, region-, window-, and project-adjusted probabilities. Primary discovery is optional; fixed-region confirmation is on by default. The WPGMA context is built once per affected cyclic round, the random prefix survives round invalidation, and unchanged triplets also benefit from the shared XOverList-style shortlist.
  • Adaptive worker batches targeting roughly 40 ms, reusable scan buffers, progress serialization/posting/rendering limited to once every 100 ms, exact monotonic phase/round timing, a visible indeterminate meter while a fresh cyclic round starts, and method-aware plot downsampling that forces both breakpoints, the selected method peak, and applicable profile maxima into the browser payload. CHIMAERA displays only its selected target/parent-one trace; GENECONV displays a three-colour -log10(raw KA P) inner/outer fragment envelope; 3SEQ displays all three signed target-specific random walks on a zero-aware axis; SISCAN displays three signed, pair-associated curves selected from the strongest eligible partition/summed Z category at each window. Later-round reconstructions are labelled when the exact historical erased/fragment profile was not serialized.
  • Graceful batch-boundary stopping during primary or cyclic discovery. An unfinished round's transient signals are discarded, previously completed events are retained, ordinary evidence finalization runs, and Review receives a normal result marked user-stopped.
  • Combined strongest-first cyclic detection: scan all eligible triplets in the supplied RDP/GENECONV/BootScan/MaxChi/CHIMAERA/SISCAN/3SEQ method-major order, reconcile the best event, infer its co-recombinant group, erase the event tract, then retain XOverList-style summaries for unchanged exact working triplets. Triplets that touch erased rows or new fragments run fresh kernels; unchanged signal-bearing triplets replay their compact summaries. An unchanged triplet that was clean on its first full screen never runs a discovery kernel again. The one-time 3SEQ post-erasure split refresh applies only to a previously signal-bearing triplet, so it cannot defeat that clean-negative rule. Committed events are durable action-cache entries; automatic erasure uses the selected role's final ConsensusOK/FinalTrim distance list, and a 64-round ceiling preserves completed events and proceeds to Review with an explicit cycle-limit-reached result. For changed rows, the supplied sparse DoPairs rule runs a method only when all three original-sequence pairs were enabled by an XOverList signal touching the selected RList; this replaces an unnecessary all-dirty-triplet rescan without changing the retained unchanged-signal shortlist.
  • The source MakeMCCorrection factor is fixed from the initial scan plan while each later round's actual fragment-expanded workload remains visible separately. GENECONV review distinguishes raw GCCalcPValP2 probability from the RDP5 XOverList-equivalent project-corrected value and makes clear that later BURT polishing does not recalculate either value.
  • Event review and CSV export explicitly flag MaxChi+CHIMAERA-only support because the manual treats those methods as closely related rather than independent confirmation.
  • Source-shaped fragment re-entry below the supplied 100,000-site cutoff. Erased tracts become gap-padded working sequences with original/event provenance; same-origin copies cannot share a triplet, short/duplicate fragments are omitted, and a visible 256-fragment cap bounds browser cost. A new fragment receives one complete follow-up pass; if it participates in no signal it is removed before another event adds rows, using supplied DropSeqs swap-with-last compaction and exact remapping of every live shortlist/signal working index.
  • Event hypotheses use the supplied detectable-signal rule: two shared original sequence identities and greater than 30% symmetric tract overlap, including fragment-assisted signals.
  • Follow-up RDP profile checks of every masked (but not disabled) sequence, retaining trace-like profiles separately from statistically significant signals.
  • Three role hypotheses per event, with each anchor sequence treated in turn as the presumed recombinant as required by the RDP5 reconciliation workflow.
  • The supplied distance-pattern correlation layer: source-defined breakpoint flanks, five region profiles, three paired six-value Pearson tests, five inverse-category relabellings, dominant-pattern and distance-triangle warnings, MakeGoodC overlap eligibility, the active MakeACOR affinity gate, exact MakeRList dual-correlation override, positive aggregate scoring, StripDupInv, and the active first FinalTrim duplicate-correlation cleanup.
  • The source FinalTrim OKSeq 6 nearest-nonrecombinant fixed-point pass, including post- StripDupInv swap-last list order, exact 0.83/0.95/0.99 gates, direct-event availability, outside/inside tree bounds, and paired breakpoint-distance veto. Its retained lists feed both ascending final expansions and selected-role pruning.
  • The complete active FinalTrim matrix-score family for OKSeq 7–14: collapsed/raw tree position, whole-tract and breakpoint JC distances, explicit source-zero slots 10/11, and FindActualEvents/MakeMatchMatX2P detected-region distance for 14. Warning gates, asymmetric ties, saturation, penalties, closest-pair modifiers, the active bare-CompMat index quirk, and the raw subtotal feed the active completed consensus score and remain auditable.
  • The supplied CalcMatchY evidence path for OKSeq 17 and 18: four bounded 40-variable-site flank walks, VB half-to-even window rounding, signed match states, circular rolling smoothing, regional-product score, six breakpoint checkpoints, standard threshold class, and the opening ConsensusOK raw-tree topology-consistency filter. Raw/filtered classes and bounded fallback state remain auditable; available rows now drive active grouping.
  • The active RFF=0 FinalTrim completion: nearest-nonrecombinant thresholds are retained for the correlation-gated expansion, raw whole-tree parent bounds drive the second expansion, and the selected-role branch preserves swap-last deletion plus the source's inherited third-list index.
  • Completed ConsensusOK scoring and membership: OKSeq 0–6 and 15, CheckPatternX, RCorrX, the source's declared-Long NS narrowing, topology-filtered OKSeq 18, primary thresholds, exact-distance equivalence widening, straggler collection, and all-role empty fallback now rebuild each distance list before the manual's two-of-three group is formed.
  • The shared selected-role conservative cleanup after the RFF guard: direct-distance outlier ranks, raw/direct topology constraints, swap-last removal, bounded-cluster admission, and the source's always-true x = x strict-inlier branch now finalize each distance list.
  • The primary-RDP post-group recheck: every finalized nonrepresentative list candidate is rerun against its role's two representatives with the supplied LowP * 100000 lift. Emitted signals, candidate-recombinant signals, event-overlapping traces, ordinary corrected significance, and the best tract/probabilities remain separately auditable.
  • A separate source-shaped MaxChi confirmation kernel from the supplied FastRecCheckMC2 path. It builds the three variable-site pair profiles, applies the DLL match-difference screen and ChiPVal2P approximation, excludes native MissingData/prior-erasure and linear-edge windows, grows the strongest peak, and retains raw, within-triplet, and project-corrected probabilities separately. The event representative triplet and every finalized nonrepresentative distance-list row are rechecked. The scan uses rolling match totals, so its strongest-peak pass is linear in variable sites rather than linear in both sites and window width.
  • Source-shaped CHIMAERA, GENECONV, SISCAN, and 3SEQ late corroboration over those same representative and finalized-list triplets. FastRecCheckChim rotates all three targets, ordinary GCXoverD retains its best six-track KA fragment, and TSXOver(1) evaluates both split walk orientations plus the supplied inverse-parent/inverse-interval list copy. SISCAN retains the nearest WPGMA outlier and full fixed-region vertical-permutation probability chain. These records keep their own probability scopes and never move a reconciled event.
  • Six Jukes–Cantor event trees using the supplied single-precision Clearcut NJ path, Microsoft-CRT SEQBOOT2 ten-replicate weights, retained-base-tree support pseudocount, VB-rounded percentage, TreeMidP/UltraTreeDistP midpoint rooting, 50% node collapse, and MakeTreeArrayXP2-shaped rank-coded topology distances. The raw/collapsed rank matrices now feed TreePhPr, TreeSubPhPr, TreeSubDist, TrpScore, paired-tree membership, and downstream role consensus; source-serialized five-decimal branches remain separate display data.
  • Iterative detectable-set closure and the manual's complete “present in at least two of three” co-recombinant group for every presumed-recombinant role.
  • An auditable source-decision-tree role recommendation covering direct, raw-tree and collapsed-tree PhylPro families, leave-one-role-out and displacement scores, RCompat, OU/list context, VisRD dMax, SimpleDist, and the currently mapped combined prizes. Native weight parity remains explicitly false while the traced RList/ListCorr dependencies are still being reconciled.
  • Ordered event review with accept/reject decisions, editable roles, breakpoints, and complete co-recombinant membership. The current and automatic groups remain separate and auditable.
  • An on-demand graphical breakpoint inspector derived from the original alignment. It prioritizes the recombinant, both parents, current/automatic co-groups, masked traces, and supporting evidence; shows bounded windows at both breakpoints; handles circular origin wrapping; and colour-codes parent matches without transferring the full alignment to the main thread. Persistent event context applies the manual's RDP-specific deleted-tract rule: a boundary within one RDP window, counted in information-rich triplet positions, is marked uncertain, while immediate tract contact remains distinguishable in JSON, CSV, and review UI. Its nearest-informative-state bracket is explicitly a review aid. The separate supplied BURT/BenHMM path now performs the three-state seeded HMM training and reports signed source-labelled 99%/95% ranges, HMM positions, breakpoint movement, missing/gap adjustments, and revert state without conflating either output. The supplied default-enabled “polish breakpoints” option is available in scan settings, retained in project checkpoints, and can be disabled to preserve the primary RDP coordinates.
  • An on-demand graphical tree inspector for the same six event regions used by reconciliation. It compares whole-tract and both breakpoint pairs, labels retained-fragment leaves, and transfers compact saved edge lists rather than rebuilding trees or moving distance matrices to the interface. The active RDP 5.93 event path requests zero bootstrap replicates, so no branch-support collapse is presented for these trees.
  • An on-demand PHYLPRO event inspector mapped from supplied FindSubSeqPP, PXoverD, MakePDstMat/UpdatePDstMat, and PPRegression. It correlates left/right Hamming-distance vectors for the recombinant and both parents, supports the two supplied gap policies and optional self inclusion, and labels the active polymorphic-column map. Only those three target rows are rolled, preserving their coefficients while reducing ordinary plot work from O(LN²) to O(LN). The profile is diagnostic only: supplied RDP5 implements no PHYLPRO significance test, so it emits no p-value or discovery signal.
  • A batched erase/fragment/re-scan cycle that re-identifies later events after a correction or rejection. Core/API guards enforce review order; mid-repair project reload drops the stale tail, remaps retained signal anchors, and resumes at the changed event.
  • Reloadable project JSON (v1alpha19, accepting v1alpha1v1alpha19), expanded event-level CSV, full, enabled-only, and masked/disabled-only curation FASTA directly from the loaded dataset, accepted-group sequence removal, accepted-tract column removal, tract-masking FASTA, and event-ordered mosaic-fragment FASTA. Event-derived alignment readiness is enforced in both the interface and WASM boundary, and exports that would contain no sequences or no alignment columns fail clearly.
  • Unsaved-analysis protection around the manual's save-often loop. Finished scans and subsequent review/repair changes visibly require a fresh project checkpoint; tab exit and destructive dataset/settings replacement warn until that checkpoint is downloaded.
  • A GitHub Pages Actions workflow that installs locked JavaScript dependencies, provisions a pinned Emscripten toolchain, checks the C ABI/worker/version/schema contract, TypeScript, the linked BootScan/cache core, supplied-source SISCAN, event-tree, and PHYLPRO cores, builds the compatibility WASM target and Vite site, validates the deployment artifact, and publishes it without a separate hosting service.
  • A responsive, accessible React interface that requires no application server.

This is not yet a parity-validated RDP5 replacement. The fully exploratory and automated query-vs-reference RDP/GENECONV/BootScan/MaxChi/CHIMAERA/SISCAN/3SEQ workflows now reach an end-to-end reviewed-event result and alignment variants, including the active late list-build and selected-role cleanup paths. MaxChi and CHIMAERA exploratory discovery are active in source, but their indexing, smoothing/destruction basins, preliminary role assignment, and cross-method event order still require native golden comparison before either can be called parity validated. Native MaxChi manual-doublet/permutation, CHIMAERA permutation/full late-event-reconstruction, and GENECONV permutation/manual-pair/alternative-indel/full late-event-reconstruction modes are not yet represented. The ordinary ignored-indel KA GENECONV path is source-shaped active but unvalidated. The ordinary 3SEQ exact/random-walk, primary BootScan distance-mode, and SISCAN paths are likewise source-shaped active but unvalidated. SISCAN random/most-distant outlier, alternative category/gap, and manual modes remain open. BootScan tree/similarity/permutation/manual modes and literal full edge-warning/catalogue behavior remain open. The PHYLPRO event-review profile is active and host-regression-tested but remains native-unvalidated; its explicit compact context map repairs the supplied RevSeq direction, while complete-window linear behavior is a documented adaptation. The source's encoded A/C/G/T counter reset is preserved, not treated as a defect. It never participates in exploratory discovery. Later cyclic rounds now retain the supplied FindSubSeqTS2 inclusive position map and CheckSplit3Seq/SubPVal missing-run trim, reverse-orientation retry, and corrected-P re-gate. The supplied two-orientation TSXOver(1) representative/finalized-list recheck is also active; manual permutation envelopes and full late event-catalogue reconstruction remain explicit boundaries. Source-shaped FastRecCheckChim strongest-target, ordinary-kernel GENECONV, and 3SEQ Findall rechecks remain non-coordinate-changing and require native golden comparison. Unported role-method families and their post-group method-stack rechecks, broader statistical breakpoint probability/confidence diagnostics, and golden native comparison are still open. The manual's RDP deleted-tract one-window uncertainty rule and the primary-RDP post-group signal/probability recheck are active, as is the distinct BURT/BenHMM statistical 99%/95% confidence and repositioning path; other detection and breakpoint-probability families remain later milestones. Exports keep these fidelity boundaries explicit and do not label the implemented native-weight subset as full desktop method parity.

See STATUS.md, docs/fidelity-notes.md, the late-consensus source trace, and the breakpoint-uncertainty source trace, the BURT/confidence source trace, and the MaxChi discovery trace, the MaxChi recheck source trace, and the CHIMAERA discovery trace, the GENECONV discovery trace, the 3SEQ discovery trace, the cyclic-shortlist trace, and the BootScan discovery trace, and the SISCAN source trace, and the event-tree kernel trace, and the PHYLPRO review trace, and the Session 26 handoff before interpreting results or starting the next phase.

Deploy with GitHub Pages Actions

The repository includes .github/workflows/deploy-pages.yml. To publish it:

  1. Put the contents of this rdp-wasm directory at the root of a GitHub repository.
  2. In GitHub, open Settings → Pages and set Source to GitHub Actions.
  3. Push the workflow to the repository's default branch. The same workflow can also be started manually from the Actions tab.

Only a default-branch push deploys automatically; pushes to other branches create a skipped run. The workflow publishes dist/ only after the expected HTML, Emscripten loader, and valid WASM binary are present and every built URL is safe for a Pages project subdirectory. The existing relative Vite base also supports a user/organization root site or custom domain.

The workflow builds both the pthread and single-worker modules. Because GitHub Pages cannot set COOP/COEP response headers directly, a small same-origin service worker adds them and performs one automatic reload on first use. Browsers that permit this become cross-origin isolated and select the threaded module; blocked/unsupported service workers transparently retain the single-worker module. No sequence or analysis data are cached by that service worker.

Automated query vs reference workflow

  1. Load the aligned dataset. Names using the manual's REF-A<sequence name> convention are detected into editable reference groups; “Detect REF names” can reapply that mapping.
  2. In the dataset table, leave a role blank for a query or enter a positive integer for its reference group. Select individual rows—or every row matching the current name filter—to assign a group or make them queries in one action. “Compact groups” renumbers IDs by first appearance without changing group membership. Records in the same reference group are not paired with one another. Role assignment does not silently change enabled, masked, or disabled curation.
  3. On the settings page, choose Query vs reference. The plan must contain at least one eligible query and eligible references from two groups. The screen shows the exact scheduled record triplets and the separately capped group-pair × query correction factor.
  4. Run and review in the same strongest-first cyclic order as an exploratory analysis. The role inference remains free; a reference called as recombinant is highlighted in amber and keeps its input group on the role cards, alignment view, tree view, JSON, and CSV.
  5. Download the .rdpweb.json checkpoint to preserve all assignments and review decisions. The CSV is a readable event summary, not a substitute for that complete project state.

Build and run locally

Prerequisites:

  • Node.js 20 or newer and npm
  • Emscripten with emcmake on PATH
  • CMake 3.20 or newer
npm ci
npm run build

The deployable static site will be written to dist/. Upload the contents of that directory to the static host. Vite uses relative asset paths, so a project subdirectory is supported. The host must serve .wasm files with the application/wasm media type.

For local UI development after a WASM build:

npm run dev

npm run build produces both modules. Local Vite development/preview supplies cross-origin isolation headers, while the production static bundle includes the same service-worker bootstrap used on GitHub Pages. The analysis worker selects the pthread module only when isolation and SharedArrayBuffer are available; otherwise the single-worker module remains the automatic fallback.

The generated files belong in public/wasm/:

  • rdp-core.mjs
  • rdp-core.wasm
  • rdp-core-threads.mjs, rdp-core-threads.wasm, and its generated worker helper

Project layout

Path Purpose
.github/workflows/ Locked GitHub Pages build, artifact validation, and deployment
src/ React workflow, worker client, review plots, and exports
src/workers/ Isolated WASM bridge and bounded scan scheduler
wasm/src/ Alignment readers, RDP/GENECONV/BootScan/MaxChi/CHIMAERA/SISCAN/3SEQ discovery and active rechecks, PHYLPRO review, BURT/BenHMM confidence, phylogenetics, reconciliation, trace checks, and exports
wasm/include/ Stable C ABI consumed by the worker
scripts/ Explicit WASM build entry point, host numerical regressions, and Pages artifact verifier
docs/ Workflow, fidelity, architecture, and validation handoff
package-lock.json Reproducible dependency graph used by npm ci and Actions

Data handling

No sequence is uploaded by the application. The selected file, encoded alignment, scan state, and review decisions remain in the browser tab until the user explicitly downloads an export. The project has no analytics, cookies, remote API, or persistence layer.

Source basis

This checkpoint was produced from the supplied DNA5.dll C++ source, VB6 application source, and RDP5 instruction manual. No alternate RDP implementation was consulted.

Releases

Packages

Contributors

Languages