Request-time evidence is incomplete.
A bounded HTTP policy identifies obvious automation. Longer-window behavior and entity correlation are required for activity that remains ordinary at the individual-request level.
Technical program report · Updated August 2026
Four synthetic prototypes examine the decision chain from server-side observation to controlled intervention: HTTP behavior analysis, adaptive evaluation, economic coordination investigation, and policy selection under explicit player-impact constraints.
Technical summary
A bounded HTTP policy identifies obvious automation. Longer-window behavior and entity correlation are required for activity that remains ordinary at the individual-request level.
A reference detector catches its scripted positive control but fails against bounded adaptive search. Captured failures are minimized, replayed, and retained as regressions.
In the authored scenario, accounts with distinct device and network identifiers remain individually ordinary while synchronized market behavior supplies case-level evidence.
An offline-supported intervention advances only to canary and is rolled back when shifted legitimate behavior breaches the declared friction budget.
Program objective Reduce measurable automated advantage while bounding adverse effects on legitimate users, marketplace function, review operations, and privacy.
Architecture and repository boundaries
The repositories are independently executable. The Go scorer and Red Queen harness integrate through a loopback-only observable-event contract. Connections to the economic investigation and intervention labs are architectural and schema-level; an end-to-end production pipeline is not claimed.
Observe
Bounded request-time scoring, asynchronous behavior, replay, reason codes, and analyst review.
HTTP Bot Defense LabFalsify
Equal-budget search, unseen holdouts, hard negatives, counterexample minimization, and regression promotion.
Red Queen LabInvestigate
Event-sourced exchange, sequence evidence, coordination graph, peer context, and market guardrails.
Market Integrity LabIntervene
Offline policy evaluation, evidence-support gates, player-friction budgets, canary, and rollback.
Intervention Lab| Repository | Control boundary | Implementation | Status |
|---|---|---|---|
| http-bot-defense-lab | Online HTTP evidence and progressive enforcement | Go | Integrated with Red Queen through a loopback scorer contract |
| bot-defense-red-team-lab | Adaptive search, replay, and regression generation | Python | Attacks the Go scorer and retains machine-readable evidence |
| market-integrity-lab | Offline economic and coordination investigation | Python | Independent event-sourced simulation; schema-level program link |
| marketplace-intervention-lab | Offline policy evaluation and guarded rollout | Python | Independent logged-policy simulation; schema-level program link |
Evaluation basis and metric definitions
The evaluations are deterministic or seeded synthetic scenarios authored alongside the systems. They test implementation behavior, failure handling, replayability, and evaluation discipline. They do not estimate production prevalence, precision, recall, financial loss, or causal policy impact.
A run is detected when the scorer emits a challenge or restriction. An observe decision is not counted as detection.
The relevant action- or episode-level denominator is stated beside each result. Aggregate rates must be read with persona-level intervention risk.
Synthetic realized profit under a detector divided by profit for the same strategy under a no-op detector. It is a scenario ratio, not a monetary estimate.
Account-membership measures for constructed rings, not case-level production precision and recall.
An offline policy-value estimate over mature held-out episodes. Support, uncertainty, and guardrail gates are evaluated separately.
Reported detector or evidence timing is named according to implementation. The market case threshold is retrospective offline timing, not streaming alert latency.
Results and limitations
Each result is tied to a declared synthetic fixture and a public source. Interpretation and limitations are stated adjacent to the measurement.
Project 01 · Go
Request-time scoring covers a narrower evidence window than longer-running correlation.
Constructed account recall at threshold 60
74 accounts · 2,680 requests · TP 10→24 · FP 0 in this acceptance fixture
What happened. Request-time scoring caught all ten fixed-loop accounts but missed fourteen randomized or coordinated accounts. Longer-window analysis found those constructed profiles without adding a false positive at the default threshold.
Why it matters. The online path now bounds both samples and account cardinality, defines retry and late-event semantics, and places JSONL audit delivery behind an observable fixed-capacity queue. At a lower threshold, three legitimate household accounts are flagged, making the cost of sensitivity visible.
It establishes deterministic acceptance behavior, truth separation, threshold sensitivity, and a reviewable architecture. It does not establish real precision or recall. The profiles were authored around known hypotheses and omit real-world ambiguity and drift.
Source: README · Demonstrated result and case-study evidence.
Project 02 · Python
Static positive-control performance does not measure robustness to adaptation.
What happened. Under equal 240-evaluation budgets, evolutionary search found a strategy with 23.8% more objective value than random search. The frozen strategy then avoided intervention on five unseen seeds and retained all of its synthetic harm.
Why it matters. A discovered weakness becomes an exact, 81.7%-smaller replay and a permanent regression fixture. The aim is not to claim that an optimizer beats a real service; it is to institutionalize falsification.
Macro legitimate-action friction was 0.19%, yet the power-trader persona was touched on 60% of its five seeds. The action-level average and persona-level intervention risk therefore need to be read together. Failure to discover an evasion would not prove security.
Source: committed summary.json and result interpretation.
Cross-repository engineering experiment
The Red Queen harness now attacks the actual Go scorer, not a co-located reference implementation.
Same-trace result. The v1 failure realized 2,819 synthetic credits with no score. V4 detected that exact trace, blocked 54 later market actions, and reduced realized profit to 1,200.
Operating result. A broad cadence fix was rejected after it detected every power-trader and delegated-household seed. A narrower coordination rule passed the three declared hard negatives, then fresh search found a profitable three-account evasion. The useful result is a working improvement loop, not a solved detector.
Two search seeds, three disjoint holdout seeds, 32 candidates per search method, and an authored synthetic economy. The fresh v1/v3/v4 strategies differ, so their profit values are not paired causal estimates. Only the frozen-trace replay compares identical actions.
Source: experiment report, committed machine-readable evidence, and Go implementation note.
Project 03 · Python
Repeated economic coordination can remain observable when weak identifiers differ.
Detected-ring profit · synthetic credits
−83.5% ring profit · +1.2% legitimate fill quantity · −2.6% total fill quantity
What happened. All six ring accounts used unique device, network, and session identifiers and remained below the account-only threshold. A full-batch graph/economic layer formed one six-account case with a retrospectively verified evidence threshold at T+385 seconds.
Why it matters. The paired treatment world applied a marketplace-only rule after that offline evidence threshold and exposed both reduced ring profit and liquidity tradeoffs. T+385 is not a measured streaming-alert latency.
A two-sided liquidity-provider pattern is an explicit hard negative. The graph is evidence for investigation, not proof of common ownership or malicious intent. A paired simulation is useful for mechanism testing but is not a causal estimate of a real intervention.
Source: documented synthetic scenario result and metric definitions.
Project 04 · Python
Point estimates are insufficient without support, uncertainty, and guardrail checks.
What happened. The harm-aware policy cleared overlap, effective-sample-size, maximum-weight, uncertainty, legitimate-friction, and review-load gates. The aggressive policy reported lower harm but was rejected for poor support and guardrail breaches.
Why it matters. Offline approval led only to a canary. When the synthetic population shifted toward more high-volume legitimate traders, observed friction crossed the budget and stopped the policy automatically.
The lab uses SNIPS and doubly robust estimates on a held-out period, with clustered bootstrap uncertainty. The gate now requires at least 90% global outcome maturity; the committed run has 2,697 of 2,852 holdout episodes mature (94.6%). The policy gate cannot read actor truth or unchosen outcomes; a hidden oracle audits the frozen decision afterward.
Source: committed summary.json and evaluation method.
Control and decision model
Detection and investigator surfaces cannot read synthetic actor labels. Evaluators alone join predictions to truth.
Network, device, speed, coordination, profitability, and model scores are evidence—not standalone proof.
Reason codes, thresholds, assignments, disagreement, and outcomes are explicit enough to replay and audit.
Observe, challenge, cooldown, temporary hold, and market restriction precede stronger account action.
Legitimate friction, fill quantity, liquidity, review load, and appeals belong beside abuse reduction.
Adaptive counterexamples and canary breaches are retained as regression tests, not hidden as bad demos.
| Layer | Question | Output | Failure posture |
|---|---|---|---|
| Telemetry | What happened? | Minimal observable event | Measure loss and missingness |
| Detection | What deserves attention? | Score, reason codes, evidence | Fail open where appropriate; alert |
| Investigation | What harm and relationships are supported? | Truth-blind case | Preserve uncertainty |
| Policy | What reversible action is justified? | Versioned recommendation | Reject weak support |
| Rollout | Did the action help safely? | Guardrails and outcome | Rollback on breach |
Assumptions and unknowns
The central premise is that meaningful companion-marketplace actions ultimately reach server-controlled HTTP services. This is a design hypothesis, not a statement about PlayStation architecture. Internal discovery could materially change the telemetry model, system boundaries, or appropriate controls.
Synthetic metrics transfer to production · scenario credits equal money · behavior proves shared ownership or intent · device or network identity is reliable proof · historical replay establishes causal effect · passing an offline gate authorizes deployment · failure to find an evasion proves security · any prototype reflects Sony or PlayStation architecture.
Production-validation plan
Days 0–30
Exit criterion A reviewed measurement contract and telemetry gap register.
Days 31–60
Exit criterion Supported candidates and documented failure modes—no user impact.
Days 61–90
Exit criterion One measurable intervention with governance and safe rollback.
Required discovery questions
Risk and governance requirements
Anti-bot systems can restrict account access, alter market liquidity, create support and appeal burdens, and introduce privacy risk. Detection quality is therefore only one part of the acceptance standard.
Fast or coordinated behavior can be legitimate. Shared households, accessibility tools, delegated use, and expert traders require explicit hard-negative coverage and review paths.
Evidence strength should determine action strength. Uncertain cases should favor observation, challenge, cooldown, or narrowly scoped restriction over irreversible account action.
Controls can reduce liquidity, increase spreads, delay fills, or shift abuse elsewhere. Harm prevention and market-health guardrails must be evaluated together.
Entity correlation can create sensitive relationship inferences. Collection, retention, regional use, and investigator access require explicit authorization and minimization.
A mitigation changes attacker incentives and the observed population. Permanent regressions and renewed adaptive evaluation are required after every material policy change.
Versioned decisions, reason codes, evidence provenance, exposure logging, appeal handling, and rollback ownership are required before consequential enforcement.
A candidate control should not advance beyond shadow operation without representative labels, explicit metric denominators, calibrated base-rate evaluation, hard-negative coverage, privacy review, operational capacity, reversible rollout, and named rollback and appeal owners.