AI Kanban

Add task

Todo

5
LMS: no rate limiting on public self-registration (/join/[token])Follow-up from card #218 (course join requests). The user explicitly chose to SHIP WITHOUT rate limiting, with the exposure quantified — this card tracks the accepted debt. EXPOSURE (measured, not estimated): - /join/[token] is the app's first unauthenticated WRITE path. Anyone holding a valid invite link can create accounts in a loop; username uniqueness is satisfied by incrementing a counter. - The invite token never expires (D2). The only kill switch is disableInviteToken, which breaks the link for every legitimate student too. - Cost per account: 5 documents written — user, account (password hash), session (created then discarded, since nextCookies() is not configured, so it is pure garbage), student, course_join_request. - Plus one scrypt hash at N=16384 r=16. Memory is 128*N*r = 32 MiB and ~70ms CPU for Node's NATIVE scrypt; better-auth uses the pure-JS @noble/hashes implementation, typically several times slower. This is the sharp end — a few dozen concurrent requests pin a CPU and push memory hard. - better-auth SHIPS a rate limiter but it runs in the HTTP router's onRequest hook. The server action calls auth.api.signUpEmail IN-PROCESS, so the router never runs and the limiter is bypassed entirely. proxy.ts's ALLOW_PUBLIC_SIGNUP kill-switch likewise only gates POST /api/auth/sign-up/email — this action creates accounts regardless of that setting. - Where the junk lands: StudentService.listStudents is a bare find({}).sort() with no pagination (the admin students page renders every row). UserRoleService.listUsers caps at 100 but still sorts on an unindexed email field. The join-request queue (step 27) has no pagination either, and every flooded request lands on the SAME course. RECOMMENDED MITIGATION (declined for now): a per-course cap on outstanding pending requests, enforced in joinSignupAction before registerStudent — one countDocuments, refuse above ~200 with the generic invalid-invite message. No new infrastructure, no index, and it refuses BEFORE the scrypt hash so it caps the CPU lever rather than just the row count. A time-windowed variant (requestedAt > now-1h) is strictly better but wants an index on requestedAt, which would be this repo's first (there is no index-bootstrapping mechanism today). NOT recommended: IP throttling in proxy.ts — Next middleware is stateless across instances, so it needs external state to mean anything.
4 weeks ago

In Progress

0

Staled

58
Local Claude Code context/cost readout: parser + status script + skill (+ VSCode bridge)Build a LOCAL (no network) readout of Claude Code context + cost, as one shared parser feeding three consumers. Remote sync + dashboard is a separate card. ## Deliverables (build in this order) 1. **Parser library** — reads the live session transcript, returns context size, per-turn cost, subagent cost, session total. Prototyped + validated 2026-08-01. 2. **Status script** — standalone, prints ONE line to stdout. Wire to `statusLine.command` (works in terminal today). 3. **Skill** — deep on-demand view, same parser. Works in BOTH terminal and VSCode extension today, so it unblocks the ambient-display problem. 4. **VSCode status bar bridge extension** — thin wrapper that shells out to the status script and renders stdout in the native status bar. Works in VSCode + Cursor. Build only if the skill proves insufficient. ## Verified mechanics - Transcript path: `~/.claude/projects/<cwd with / replaced by ->/<$CLAUDE_CODE_SESSION_ID>.jsonl` — verified working. - **Context size = the LAST assistant record's `input_tokens + cache_creation_input_tokens + cache_read_input_tokens`.** NOT a sum across turns (that gives cumulative usage, ~15x larger and wrong). - **CRITICAL: dedupe by `(requestId, message.id)` keeping MAX `output_tokens`.** Claude Code rewrites each assistant message to the JSONL repeatedly while streaming; early copies hold partial counts (e.g. 7 then 1505). Taking the first/last naively misreports the headline number by up to ~200x. This is the single highest-risk bug in this card. - `input_tokens` being ~0 is CORRECT — nearly all context arrives as cache_read. Do not "fix" it. - Subagent transcripts: `~/.claude/projects/<slug>/<session-id>/subagents/agent-*.jsonl` (nested one level deeper; a one-level glob misses them entirely — they were 27% of historical spend). - Pricing: cache read 0.1x input; cache write 5m 1.25x, **1h 2.0x** (`usage.cache_creation.ephemeral_{1h,5m}_input_tokens`). Sonnet 5 is on intro $2/$10 per MTok **until 2026-08-31** (list $3/$15). - Context windows: 1M for Opus 5/4.8/4.7, Sonnet 5/4.6, Fable 5; 200K Haiku 4.5. Model string may carry a `[1m]` suffix — strip before lookup. ## Status line format `◐ 43.8% (438K/1M) · turn $0.28 (+$0.11 sub) · carry $0.22/turn · ⚠ cache miss` - context % + absolute vs the model's real window - last turn cost, subagent broken out separately with `+` - **carry cost** = `context * input_price * 0.1` — what the NEXT turn costs before you type anything. The headline insight: context is a recurring bill, not a capacity gauge. On a live 438K-token session, 78% of each turn's cost was just re-reading context. - warnings: cache-prefix invalidation (large cache_creation + small cache_read => that turn cost up to 20x the cached equivalent), and context crossing a threshold ## Locked decisions - Status line REPORTS AND WARNS (not purely informational). - Subagent cost DISPLAYED, broken out separately (not silently rolled in). - Skill covers CURRENT SESSION ONLY, fully local, no network. - Scope is claude-code only — dropped the original cross-agent mirroring to cursor/antigravity (transcript format, session-id env var, and cache-tier accounting have no equivalent there). ## Performance note (VSCode bridge) Watch the project dir and read only the TAIL (~256KB) of the transcript — do not re-parse a multi-MB file on every append. Still apply max-output_tokens across copies in the tail, or the status bar flickers through partial streaming values before settling. The bridge cannot ask Claude Code which session is active (no API); infer via newest-mtime `.jsonl` in the workspace's project dir.
2 months ago
Cross-machine Claude Code usage sync + dashboard UIAggregate Claude Code usage from BOTH machines into one remote DB, with a web UI for visualisation. Sibling of card #112 (local readout) — reuses the SAME parser library; build #112 first. Motivation: ccusage is accurate but per-machine by construction, and the JSONL contains NO account/org/machine identity — so cross-machine and per-account attribution can only be solved at sync time, never reconstructed later. ## Sync behaviour (locked) - **Triggers: `SessionStart` (inline upload) AND `UserPromptSubmit` (dirty-marker only, detached).** - `SessionEnd` is REJECTED — does not fire on tab close, crash, or kill -9. - `UserPromptSubmit` must NOT do network I/O inline: it runs BEFORE the model sees the prompt, so it is in the per-turn latency path and a hung connection stalls the session. It writes a dirty marker and exits (~5ms); upload happens detached. Needs a real detach — `& disown` is not always enough since Claude Code can wait on inherited fds. - `UserPromptSubmit` structurally never syncs the last turn of a session; `SessionStart` catches it on next launch. - **No cron, no launchd, no daemon.** The JSONL on disk is the durable source of truth and survives every failure mode; the hooks are freshness, not delivery. Staleness is bounded by "when you next open Claude Code on that machine" — an idle machine generates no new usage anyway. - **On failure: do nothing, exit silently.** No error surfaced into the session, no retry logic. ## Data rules - **Aggregates ONLY — transcripts NEVER leave the machine.** The JSONL contains source code, file contents, and all tool output. Ship token counts per message. Entire history is ~46K rows. - **PK `(request_id, message_id)`** — globally unique. Makes re-sync a no-op: no watermark, no "what did I already send" state to corrupt, no harm in overlapping ranges. - **Self-healing upsert** — `on conflict (request_id, message_id) do update set output_tokens = greatest(excluded.output_tokens, usage_events.output_tokens)`. This makes the streaming-partial bug IMPOSSIBLE TO PERSIST: a row written from a partial record is corrected upward on any later sync. Build this in rather than trusting the reader. - **Stamp at write time: `machine_id`, `account_uuid`, `org_uuid`** — none exist in the JSONL. `machineID` + `oauthAccount.{accountUuid,organizationUuid}` come from `~/.claude.json`. - **Store `cost_usd` computed at sync time**, plus a separate `pricing` table for audit. Do NOT compute cost at query time from current prices: Sonnet 5 intro pricing ($2/$10 vs list $3/$15) expires **2026-08-31**, and query-time pricing would silently reprice all pre-expiry history. - Include `is_subagent` (subagent transcripts are nested at `<session-id>/subagents/agent-*.jsonl`; ~27% of historical spend). ## Backfill One-shot script per machine over existing history (~46K rows, ~$5.9K of usage). Same parser, same idempotent upsert. ## Dashboard UI Web app over the synced data. Wanted views: - cost per day / per machine / per project / per model - subagent share of spend over time - cache efficiency (read vs write ratio) — spikes indicate prefix invalidation - session drill-down - optional: join `~/.claude/usage-data/session-meta/*.json` (tool counts, git commits, lines added/removed, interruptions, response latency) for cost-per-commit style metrics. NOTE: only covers 134 of 175 sessions and looks like a one-off snapshot from 2026-07-01 — verify it still regenerates before depending on it. ## Open (deferred by user until behaviour was settled) 1. **DB host** — recommend a NEW dedicated Supabase project (Supabase MCP already wired; Postgres over REST means the sync is a plain curl with no client library on either machine; RLS keeps it private). Alternatives: existing mainnet/testnet project (mixes personal telemetry into something else), or Neon via personal-infra Pulumi (consistent with existing IaC, but needs a PR/deploy cycle and free tier has a 6h retention cap). 2. **Account topology** — do both machines run under the same Anthropic account (`b7bc187e-7b55-4c62-8977-c069bffdfd83`), or is one a work account? Decides whether `account_uuid` is a constant or a real schema dimension. Recommend modelling it as a first-class column regardless, so adding a work account later needs no migration.
2 months ago
Fix near-bankrupt sleeve inflating the 1/N blended book (fake +9193% bar)Run #28 showed a single +9193% bar (equity 5,254 -> 488,352 in one hour). Root cause found: on 2021-04-16 22:00 dogeusdt's fold equity fell 2,045 -> $1.62 (Doge squeeze over a naked short), then recovered to $1,329 = +82,002% in one bar; _blend_1n averages the 9 coins' RETURNS index-wise so that became +91.1 blended. TWO compounding causes: (1) the S4 bankruptcy floor is a knife-edge at exactly equity<=0, so a sleeve at $1.62 of $10,000 is economically dead but treated as fully alive with astronomically leveraged percentage returns; (2) _blend_1n averages percentages, implicitly re-funding every sleeve to a full 1/N each bar, so a dead sleeve's +82,000% is credited as if earned on real capital. DAMAGE: run #28 reads +409% total but that ONE bar is 92.9x — everything else is -94.5%; mean-of-fold Sharpe was -1.105 and the DSR gate still PASSED at 0.970. Run #27 (the z-ladder research run behind the "scaling-in beats single-entry, +0.366 stitched" conclusion) has the SAME artifact from the same Doge event (+1292% bar = 13.9x; everything else -87.2%), so that headline is NOT supported and the project memory + docs/features/zscore-ladder-mean-reversion/spec.md are now wrong. FIX (user approved both): (1) insolvency threshold instead of a knife-edge — liquidate and hold dead below a fraction of starting capital; (2) blend DOLLARS not percentages so a dead sleeve contributes its actual negligible capital. Then re-run #27/#28 configs for honest numbers and correct the memory + spec.
2 months ago
Template-clone the test database fixture instead of rebuilding per testBranch perf/test-db-template, stacked on fix/anvil-fixture-flakes (PR #717). Follow-up to card #109: the dbFixture hook timeout was MASKED there (15s->30s), not fixed. Root cause is that testDbFixture runs with each:true and each run does initdb + boot a whole postgres server + replay ~3116 lines of schema DDL + seed mock data - measured 3114ms - for EVERY test, with 4 jest workers doing that disk-heavy work concurrently. That contention is what blew the hook in a run whose lint/compile times were completely normal. Measured with a scratch probe: full build 3114ms vs `create database ... template` clone 242/317/349ms (mean 303ms) = 10.3x on database provisioning alone. Approach: hide the whole mechanism inside lib/core so dbFixture's returned shape (adminClient, pgInstance, db, readOnlyDb - used in ~96 places across the suites) is unchanged. makePostgresInstance now lazily builds ONE server + template database per (worker, preamble) and clones per call; PostgresTestInstance.kill() drops the clone rather than killing the server. New shutdownPostgresTemplates() export, called from an afterAll registered inside testDbFixture, so template lifetime is per test FILE - matching how anvil and the seed data already work. Rejected alternatives, with reasons: each:false (trades away the per-test isolation ~660 tests are written against; marketRoute alone asserts an exact count of closed markets, so cross-test pollution would surface as order-dependent failures - a nastier flake class than the one being fixed). Schema-per-test (user's own objection, correct: it still replays all the DDL into the new schema and re-seeds; only saves initdb+boot. Template skips the DDL entirely because it is a file-level copy).
2 months ago
review-changes: auto-tier the pipeline by diff size to cut token costThe claude-code review-changes skill spawns 10-12 sub-agents per review (1 holistic + up to 6 lenses + 2-4 verifiers + 1 merge). Each sub-agent pays ~30k input tokens of fixed overhead (system prompt + tool schemas ~12k, injected repo CLAUDE.md + .claude/rules/* ~12-15k, node file + lens-common.md ~2.5-4.5k, HOLISTIC.md ~2k) BEFORE reading any diff => ~300-360k tokens of pure overhead per run. Holistic gates lenses on applicability (does the diff touch auth / a migration) but never on SIZE, so a 40-line diff costs about the same as a 3000-line one. DECISION (user, 2026-08-07): auto-tier inside the existing skill (no separate lite skill, no manual override arg), and group the lenses on the cheap path. Tier ladder, decided by holistic (which already emits an eligibility verdict): - stop / single-inline-pass — unchanged, 1 agent - compact (small diff, neither security nor architecture fired) — 1 grouped reviewer covering all applicable lenses, capped verify, sonnet merge => ~3-4 agents - grouped (medium, or small+risky) — 2 agents: `mechanical` (correctness+quality+tests+performance, sonnet) and `deep` (security+architecture, session default) => ~5-6 agents - fan-out (large: >~25 files or >~1000 changed lines) — today's 6-lens pipeline, unchanged Security/architecture never share an agent with the mechanical lenses when they fire, preserving the skill's existing "never discount security" rule. Scope: claude-code variant (skills/claude-code/review-changes/). Cursor variant shares the node files and may be ported; antigravity is a single-agent inline design with no fan-out, so the tier ladder does not apply the same way.
2 months ago
Balance BE jest CI shards by measured duration instead of file-count hashDistinct follow-up to card #136 (template-clone fixture, PR #725) — same goal of cutting BE CI time, different mechanism. upredict-backend CI splits ~107 belief_locker test files across 4 shards using jest's DEFAULT sharder, which sorts files by SHA1 of their rootDir-relative path and takes an equal-COUNT slice. It has no knowledge of runtime. Reproduced the hash assignment in Python: 107/107 files matched the observed CI split exactly. Measured on run 31163619775 (branch perf/test-db-template): per-shard work is 1742 / 1366 / 1907 / 1802 seconds — a 40% spread. Shard 3 is the critical path at 521.9s of jest, 647s of job wall-clock. Second, independent defect: within a shard, jest's sort() falls back to FILE SIZE descending when no timing cache exists (always true on a fresh CI runner). On shard 3 that dispatched serverNoBets.referrals.test.ts (3.8KB but 86.4s) last, at t=435.5s, extending the shard from ~436s to 521.9s. A 4-worker list-scheduling simulator driven by file-size order reproduces all four observed shard wall-clocks to within 2.8s (shard 3 to 0.1s), so the mechanism is confirmed rather than inferred. Projected: duration-balanced shards + duration-desc ordering gives a 427.5s critical path vs 521.8s today — 94s / 18.1% — within 1.4s of the theoretical floor (total work 6817s / 16 lanes = 426.1s). Hard ceiling: server.reveals.advanced.test.ts alone is 403.8s, so no shard or worker count beyond 4x4 buys anything until that one file is split.
2 months ago
RISE-15509 — Allow creating associations between new ORG_TYPES from SMWiden RSC's write API so SM can push org + relationship changes for the new SM-aligned org types. Epic RISE-15475, implements PRP-2222. Sibling ticket RISE-14815 (dup epic RISE-14821) proposes the same endpoint with fuller AC — confirm which is live. SCOPE (clarified by operator, not in the ticket): - RSC OWNS association + org-type data. Not mirroring SM's schema, not reading associations from CDC. - CDC read-side work (RISE-14881 etc.) is a separate track. - RISE-15515 (Kafka) is the same payloads async, deferred — the REST contract designed now is what that consumer reuses. - Write path is DF-gated; reject when DF off. - RSC adopts SM's org-type names. Direction for symmetric pairs: SM decides, RSC stores as sent. THE CORE MISMATCH: SM stores org type once per ecosystem (ecosystem_organization.type). RSC stores it once globally (organizations.type). RSC has no per-ecosystem type slot. BLOCKERS TODAY: - RSC validator rejects most pairs: partnerId in {partner,inspection_service,agency}, factoryId in {factory,partner}, ecosystemId must be retailer (validators.js:240-246). - SM drops most pairs before sending: should_process_rise_sync requires exactly {SUPPLIER,FACTORY}, silently (rise_association_trigger.py:11-17). - Adding org types touches PG enum organization_types used by 3 columns (organizations.type, onboarding_organizations.type, plans.org_type) AND needs plans rows seeded or subscription creation fails. Design doc (309 lines, all PRE measurements): /Users/quan.vo/Documents/git-repos/inspectorio/RISE-15509_SM_RSC_ORG_MODEL.md
rs-backend·feature/RISE-15509/allow-associations-new-org-types2 months ago
CCP environments phase 2: GCS zipball storage, 10MB streaming cap, build-queue + claim, fire-and-forget auditFollow-on to card #147 / PR #114 (ConcreteEngine/ccp, branch feat/environments-async-build, worktree ccp-environments-async-build/). PR #114 shipped POST validation + HTTPS-only GitHub URL check + PUT /:envKey build reports. This card covers the scope that arrived afterwards. SIX PIECES OF WORK 1. 10 MB cap on the zipball, enforced WHILE STREAMING to disk (Readable.fromWeb + byte counter + AbortController). Verified empirically: GitHub zipball 302s to codeload and sends NO content-length for real repos (express/react/linux); only a toy repo did. Repo metadata `size` is not a substitute — express reports 9843 KB but its HEAD zipball is 220,607 bytes (46x over) because size covers full git history. 2. Store the zipball in GCS (2sync bucket, alongside models/user.js createUserDirectory convention); keep only the object path on the Mongo doc. 3. GET /environments/:envKey/zipball — separate URL from build-queue so the queue payload stays small. Must stream actual bytes through CCP (NOT a signed GCS URL) because the rig can only reach CCP. 4. GET /environments/build-queue — read-only LIST of queued environments (mirrors GET /jobs/queued), employee-gated. Plus a separate POST claim endpoint that transitions queued -> building, conditionally so a second rig gets 409. 5. Unzip + Dockerfile audit (models/docker.js DockerfileAuditor is currently dead outside tests). Failure is RECORDED as a status + buildNotes, never a failed request. No unzip lib in package.json yet. 6. Delete the GCS object on the terminal PUT (success or failure). DECIDED - Zipball storage: GCS. - Audit: fire-and-forget after the POST response. Operator chose this over inline-non-fatal after being told the work is lost on restart and errors have nowhere to surface in the response. Mitigation: outcomes land on the environment doc as buildStatus + buildNotes. - Cleanup: on terminal PUT only, so a rig can re-fetch after a failed unzip. - build-queue returns a LIST, not one item. - Claim is an explicit POST by ICP, never a mutation on GET — ICP can die between read and build. - Lease / reset for environments stuck in Building: DEFERRED to a later ticket. - Status vocabulary: keep lowercase and add statuses as needed; revisit casing after development. - origin/docker-auditor-refactor: leave alone entirely. NOTE it cannot merge as-is (drops 'inprogress' while the route still writes it -> every create 500s). Stray tracked environments/test-env-1772721276469/Dockerfile stays. - Rig uses the zipball, not git clone — it has no network access beyond CCP. shell-client PR #25's clone_env_repo must be replaced, and repoToken should NOT appear in the build-queue payload. CONTEXT - CCP has NO background job mechanism (no cron/scheduler/worker) — confirmed by grep. - Base.upsert appends a full doc copy to log[] on every write, which is why bytes cannot live on the Mongo doc. - Precedent: GET /jobs/queued (routes/jobs.js:112) + Job.getQueued (models/job.js:83), employee-gated via auth.email.endsWith('@concreteengine.com'). Note jobs uses 401 there while our PUT uses 403. - ICP proxy (icp PR #37, middleware/environments.js) pipes CCP response bodies through via Readable.fromWeb().pipe(res), so binary streaming works end to end unchanged.
ConcreteEngine/ccp·feat/environments-async-build2 months ago
Fix CI hangs from apt/Ubuntu-mirror stall in Setup PostgreSQL (upredict-infra + upredict-backend)Root cause (confirmed): the `Setup PostgreSQL` step (`tj-actions/install-postgresql@v3`) runs `apt-get update`. Since 2026-08-18 `azure.archive.ubuntu.com` is intermittently unreachable from GitHub-hosted runners; apt returns `Ign:` ~25x with backoff, falls back to `https://archive.ubuntu.com`, then stalls indefinitely (no apt timeout configured). Proof: run 32227794374, shard 1 = 28s (azure mirror `Hit:`) vs shards 2+3 = 45m (azure `Ign:` -> fallback -> wedge), 27 seconds apart on the same run. Not a repo change: build.yml untouched since 2026-08-03, release-workflow.yml since 2026-07-19. Hang duration == each job's timeout ceiling, which is why infra was far worse than backend: backend python job `timeout-minutes: 10` -> 591-603s hangs; backend ts shards `timeout-minutes: 45` -> 43-45m hangs; infra release-workflow had NO `timeout-minutes` -> ran to GitHub's 6-hour default (run 32161546065 = 360 min). DONE: cancelled 4 wedged infra runs (32216545766 stuck 3h45m, 32225834718, 32230013499, 32231468988). Added `timeout-minutes: 10` + one-line comment to the 5 runs-on jobs (release-workflow.yml LintTypescript + PostgresMigrationCheck, validate_configs.yml validate-config, terraform_apply.yml terraform, common-workflow.yml config-images). `timeout-minutes` is illegal on a reusable-workflow caller, so the cap goes on runs-on jobs only; every `uses:` chain terminates in one of those 5, so coverage is complete. Rebased deploy-20260819.2-to-production onto origin/main (bd755c8, clean) and force-pushed with lease; commit 4523cc6 on PR https://github.com/SportsFI-UBet/upredict-infra/pull/720. Verified run 32233433516 progressing normally. BLOCKER on the requested follow-up: a `pgvector/pgvector:pg15` **service** container will not work in either repo. `upredict-infra/ci/scripts/typescript/src/postgresDiff.ts` and `upredict-backend/typescript/lib/core/src/postgresTestInstance.ts` both call `pg_config | grep BINDIR`, then `initdb <tmpdir> --auth=trust`, then `spawn(postgres, -p <randomPort>)` — they need postgres BINARIES on the runner and deliberately use a random port per instance so matrix targets / test shards run in parallel. A service container supplies a networked server on a fixed port and no binaries. Also note `terraform_apply.yml` has the same apt exposure via `sudo apt-get install colorized-logs`, and `Install pgvector` shells the pgdg apt script in both repos.
2 months ago
UBET-4282 In-App Currency Betting — currency model design (brainstorm)Epic UBET-4282 "In-App Currency Betting" / task UBET-4270 "Review betting code & planning". Both Jira issues are EMPTY (no description, no comments) — every requirement is derived from code or from the PO conversation. Settled: betting with in-app currency runs OFF-CHAIN in Postgres. On-chain gas is currently paid by the platform; shifting it to users needs the embedded wallet, which is not ready. Off-chain removes the cost entirely. Key code findings: payout math is ALREADY off-chain (get_bet_payout_calculations_from_result, sql_schemas/upredict_index_view_function.sql:1702) so the whole settlement layer is reusable; betting is parimutuel so the platform has no book risk — liability is only at the mint. market_bet has NO user_id (bets join to users by EVM address) and its commitment/salt/nonce/deadline-block columns are all NOT NULL, so off-chain bets need schema work regardless of the currency answer. point_ledger already carries negative rows (BetVolumeReversal), has unique(source_type, source_id) for free idempotency, but points is double precision and the display clamps greatest(0, ...), which would hide an overdraft rather than prevent it. Structured brainstorm docs in workspace/tmp/UBET-4282/ (brainstorm-in-app-currency.md = Level 0, level-1-currency-model.md = Level 1). Level 2 NOT started — needs explicit approval. PO answers so far: BP may be bet and lost, and losing rank that way is intended; play money for now; a BP bet earns no BP; betting currency is meant to run out and engagement refills it; one earning mechanism only (no separate faucet).
last month
Brainstorm: Explanation docs for Concrete Engine (AI-updatable, Diátaxis)Structured brainstorm for the "explanation" quadrant of a 4-kind doc set (specs / how-to / tutorial / explanation) covering the Concrete Engine platform. Requirement: docs must be updatable by AI when code changes. Workspace: tmp/system-explanation-docs/ — brainstorm-explanation-docs.md (index), level-0-widest-view.md, level-1-structure-alternatives.md. User decisions: audience = new engineers onboarding; location = all at concrete_engine/ root; AI update trigger = deferred ("maybe CI, later"), so design structure for it but don't build it; existing README.md + 7 repo_knowledge/ folders stay untouched (new layer alongside). Level 0 settled: purpose is onboarding-first with decision-record as the distinguishing content; the requirement is verify + partially-update + protect reasoning, NOT regenerate (explanation is the one Diátaxis kind not derivable from code); the defensible niche vs existing docs is cross-cutting "why"; intent is the doc spine with intent-vs-actual divergences marked inline (research/job-dispatch-gaps.md says the job path is broken end to end). Level 1 settled: structure = layered hybrid (00-orientation.md narrative → concepts/ → decisions/), chosen because the three layers have different volatility, which is the seam that makes partial AI updates safe. Rejected per-service mirroring since repo_knowledge/ already occupies it. Updatability = provenance anchors in frontmatter + in-file derived/reasoned zones + deferred .doc-state.yml SHA snapshot, yielding a 5-step update procedure whose step 4 is "flag reasoned-zone conflicts for a human, never rewrite". Key constraint discovered: 7 independent git repos under a non-git root, so there is no system-wide diff — docs must carry their own provenance.
last month
OpenClaw health: Chat watcher blocked by tool cap + 4 latent breakagesInvestigation triggered by the "Google Chat watcher couldn't run — no shell/exec tool" failure reported from an OpenClaw cron session. # Root cause of the reported failure — two independent walls **Wall A (the actual blocker): job ea9f9707 has no `exec` in its stored tool cap.** Created 2026-09-03 from the Telegram conversation; snapshotted a 42-tool default allowlist with `toolsAllowIsDefault: true` that omits all of `group:runtime` (exec/process/code_execution) and `group:fs` (read/write/edit/apply_patch). It keeps dir_list/file_fetch/file_write, which are NODE file ops, not host shell. `openclaw doctor` names this job and explicitly REFUSES to widen the cap ("doctor will not silently widen or rewrite it") — so `doctor --fix` does not repair it. Supported path is `openclaw automations edit <id> --tools <list>`. **Wall B: the node fallback is dead too, for a different reason.** Agent fell back to the `nodes` tool and called `dir.list`; gateway log 11:00:17 and 15:03:56 both: `INVALID_REQUEST node command not allowed: "dir.list" is not in the allowlist for platform "macOS 26.6.2"`. Live node runs OpenClaw.app 2026.8.1, which advertises `fs.listDir` — renamed; gateway (npm CLI) is 2026.9.1, a version ahead. Even `system.run` would fail: node `bins` = claude,git,gh,gog,node,curl,python3,jq — no bash/sh. Evidence it never once worked: `.chat-state/snapshot.txt` still holds the 01:14 seed from `init`; no poll has advanced it. # Other findings (unrelated to the reported error) - **Anthropic auth profile vanished.** `auth.profiles` is `{}` in openclaw.json; all four backups have `anthropic:claude-cli` (oauth). `openclaw models auth list` → "Profiles: (none)"; `secret_store_entries` empty. Breaks `skill-collection-review-main` every run and the remote model catalog refresh. Agent turns survive only because models are pinned to `agentRuntime: claude-cli` (Claude Code's own login). - **Telegram delivery dropping cron output.** Calendar digest ran ok 07:43 but all 4 retries failed `Network request for 'sendMessage' failed!`. Calendar new-meeting watch + Birthday notifier also `not-delivered`. - **Version skew:** OpenClaw.app 2026.8.1 vs npm CLI/gateway 2026.9.1 — the direct cause of the dir.list/fs.listDir mismatch. - **23:00 evening briefing** errored with "Legacy workspace setup state requires migration" — but that was its 00:58 run during the crash-loop window; the 01:08 `doctor --fix` resolved it and the 08:45 morning briefing ran clean. Expected to self-heal; watch tonight. - Minor: 1 dead-lettered telegram ingress event; no backup ever recorded; gateway service PATH missing /opt/homebrew/opt/node/bin; plaintext gateway.auth.token + telegram botToken in openclaw.json. # Done this session - Truncated `~/Library/Logs/openclaw/gateway.log` (830 MB → 0; last 2 MB kept in session scratchpad). This is the launchd StandardOutPath, which has NO rotation — OpenClaw's own 100 MB `logging.maxFileBytes` rotation does not cover stdout capture, so it will regrow. - Fixed the watchdog install mechanism. `openclaw-ops/watchdog/install.sh` symlinked the script into ~/Documents, which macOS TCC blocks for launchd's child — `watchdog.err.log` was full of "Operation not permitted" and the watchdog silently never fired until someone hand-copied the file at 01:10. install.sh now `install`s real copies of both script and plist, with the WHY in a comment. Re-ran it: watchdog live (wedge-count 0). **UNCOMMITTED** on openclaw-ops master. # Open 1. Repair job ea9f9707 (needs a decision: command payload vs agentTurn + exec). 2. Restore the anthropic auth profile. 3. Update OpenClaw.app to 2026.9.1 to clear the node command skew. 4. Commit the install.sh change. 5. Consider a durable rotation for gateway.log (newsyslog or periodic truncate).
4 weeks ago
Cross-project visual UI QA: screenshot every state, agent reviews for layout defectsTrigger: LMS `AddQuestionForm` renders its status `<output>` and error `<div>` as direct children of the sidebar/form FLEX ROW instead of inside the form panel, so the success banner becomes a third flex column that stretches full height and squeezes the form into a narrow strip. Same bug duplicated in add-pool-question-form.tsx. Unit tests passed (they assert message text, not placement); e2e passed (behavior, not appearance); lint/tsc cannot see it. Generalized failure class: the state exists, is reachable, is functionally correct, and is visually broken. Brainstorm workspace: AI-rules-repo/tmp/visual-qa-agent/ (brainstorm-visual-qa-agent.md + answers.md + prior-art.md + zoom-1/2/3). DECIDED: capture by snapping inside existing Playwright e2e specs + assertion-free tour files for gap states; judge via deterministic in-page layout probes (overflow / overlap / empty-giant / sibling-anomaly / squeeze / contrast) feeding a NAMED-DEFECT rubric, with a vision pass only on probe suspects + pixel-changed states + a small random sample; absolute judgment now, accepted shots become baselines so later runs go quiet; ship as an AI-rules skill ONLY (no npm package); emit a human contact sheet alongside the report as the zero-cost fallback. PRIOR ART: nothing does this end-to-end. github/awesome-copilot@web-design-reviewer (13.3K installs) has a strong ~60-item visual checklist worth harvesting, but its workflow is live/interactive with no archive, no batch mode, no measurement layer. software-mansion/argent@argent-screenshot-diff (14K) is baseline diffing — the shape of phase 2, unusable today with no baselines. KEY INSIGHT: a vision model asked "does this look right?" is unreliable. It becomes reliable when the question turns factual — ship layout measurements alongside the pixels, give a closed rubric of named defects rather than an open prompt, and require every finding to name an element or be dropped. The trigger bug then reads as "`<output>` is the 3rd flex child of a row, 42% width, 4% ink coverage" instead of "looks weird".
4 weeks ago
RISE-15510 — Request Bulk Assessment: filter Assessed Org / Selected Partner by standard's org type & statusOrchestrated-feature-dev run for RISE-15510 (epic RISE-15475). The BULK half of the org type/status filter that RISE-15478 shipped for the SINGLE request — 15478 explicitly scoped bulk out (D37: "the Activation List tab ships unfiltered... Operator declined the fix at Phase 5b triage. Consequence: TC-016 does NOT pass"). This ticket closes that bypass. AC1: selecting an Activation list for Assessed Organization lists only matching orgs, shows a warning with matched/excluded counts, and offers a "Download log file" CSV (Organization ID | Name | Type | Status) of the excluded ones. AC2: the per-row Selected Partner dropdown lists only matching orgs. AC3: if the Admin narrows the settings while the form is open, clicking Request auto-removes the now-invalid orgs and shows an alert. Prior art already on master: stakeholder_org_restriction.service.js (buildStakeholderOrgRestrictionPredicate, 4 slots), organizations.sm_type/sm_status (RISE-15577), the RISE-15673 ecosystem-owner status exemption. Unmerged prior art on the RISE-15592 branches: BE validateStakeholderOrgRestrictions aggregate validator, FE orgQualificationRule.ts + InvalidStakeholderOrgAlert. Workspace: tmp/RISE-15510/ (TICKET.md, CONTEXT.md written). QA posted 28 manual test cases on 2026-09-06 (Jira comment 516664). KNOWN CONFLICT to resolve with the PO: RISE-15478 and RISE-15592 both settled on keep-and-flag a disqualified org, never clear it. AC3 here asks for auto-removal. The alert copy is byte-identical across all three and already says "have been removed".
rs-backend·feature/RISE-15510/bulk-request-org-type-status-filter +24 weeks ago
UBET-4297 BE: capture profile SEO/embed image once a dayTicket UBET-4297 "Profile img embedding capture once a day" (Task, Todo, reporter Daniel Jiwoong Im, assignee Quan Vo, Sprint 118, 1d estimate, epic UBET-4272 "propagation"). NO description and no comments in Jira — requirements derived from code. Running the orchestrated-feature-dev pipeline. Workspace: workspace/tmp/UBET-4297. ESTABLISHED FROM CODE (2026-09-08): - Existing pipeline is markets-only: upredict-backend/typescript/services/seo_image_generator is a Lambda that puppeteer-screenshots {originUrl}/markets/{id}/place?preview=true, waits for .preview-content, uploads JPEG q85 to S3 via presigned URL, inserts into market_seo_image. Cron cron(2/5 * * * ? *) every 5 min, lambda timeout 900s. - It is CAPTURE-ONCE-EVER: selection query guards on `not exists (market_seo_image row) and m.status = 'Open'`. Nothing in this codebase re-captures on a schedule — "once a day" is genuinely new behavior. - Public profile URL EXISTS on FE main (daa1daf): canonical /@<alias> (ROUTES.PROFILE = "/@:userAlias"), src/app/[alias]/page.tsx resolves alias->userId server-side then renders <ProfilePage userId>. Legacy /profile/<alias> 302s there. - FE profile page has NO generateMetadata (metadata = {}), and NO preview render (MarketPreviewProvider / .preview-content are market-only). Those are UBET-4273 (Nick Mai) and UBET-4301 (Nick Mai), NOT this ticket. - Backend has NO user_seo_image table — only market_seo_image (market_id, image_url, inserted_at, unique(image_url)). - Backend already serves public GET /users/:userId/profile and GET /users/resolve-id (routes/userProfile.ts). SCOPE PER USER: backend only. The generator loads the profile URL with ?preview=true appended; FE will handle that param later if changes are needed. PRIOR ART ON THIS GENERATOR: card #212 (reel SEO image captures TikTok consent dialog; notes the never-retry behavior) and card #220 (UBET-4323 reel thumbnails in SEO images).
4 weeks ago
Explanation docs: prompt evaluation + ConcreteEngine/documents repo with Gemini CIPhase 1 deliverable from tmp/system-explanation-docs/phase-plan.md: "Prompt: how to read the source repos; what belongs in the map versus in specs; house voice; the divergence marker rule." User asked to (1) write several prompt versions, (2) spawn Haiku sub-agents to generate the docs with each, (3) rate the prompts. Credential/GH-Actions setup explicitly out of scope per user. Experiment design: 3 variants forming a LADDER so the rating isolates which layer of prompt engineering earns its keep. - V1 Brief — intent only (~15 lines), trusts the model. Baseline. - V2 Contract — V1 + output skeleton, per-section content contract, explicit exclusion list, divergence-marker syntax + gap-ID table, house voice. - V3 Procedure — V2 + mandated reading order/grep targets + every claim carries file:line + a self-check pass. Fairness controls: identical harness preamble (role, cwd, output path, forbidden reads), same model (haiku), same 7 source repos. Root README.md and */repo_knowledge/** are FORBIDDEN reads for all three — a real runner only checks out source repos, and letting agents read the hand-written README would grade paraphrase rather than generation. Reading-strategy guidance is deliberately kept OUT of the preamble since that is exactly what V3 tests. Target artifact: the system map (architectural half only per D1 — what calls what + user flows; no rationale, no mechanism). Grading ground truth: root README.md + research/job-dispatch-gaps.md (gaps G1-G8, S1-S6) read into context this session. Workspace: scratchpad/prompt-eval/ (prompts/, runs/v1..v3/).
4 weeks ago
orchestrated-feature-dev: coverage-dedup gate — don't write a test for a behavior already coveredStructural gap found on the UBET-4337 run: the pipeline turns EVERY planned behavior into its own test + commit, and no phase asks "is this already covered?". Phase 3b only adds behaviors, node-bdd-step writes one test per behavior by construction, and the quality gate reviews test QUALITY but never test REDUNDANCY. Concrete failure: an 82-line integration test (commit f7c7488a, since dropped) that composed two facts each already proven elsewhere — settlePointsMarkets.treasury.test.ts:187 already asserted isSystemAccount === true, and `not au.is_system_account` in generate_leaderboard was pre-existing untouched code. User rejected it in review: "we already have tests for this, do not need to add new tests here for the sake of behaivor right?" Fix = a coverage-dedup gate that runs BEFORE the test is written (the cost avoided is the test + the commit + the review round). Test: "would this test fail if I deleted it and changed nothing else?" Distinguish "already covered" from "covered only in combination" — two CHANGED components meeting for the first time IS worth a test; one changed + one unchanged already-tested thing is NOT. A no-test behavior must be explicit in IMPLEMENTATION_PROGRESS.md naming the covering test, and gets no commit. Related: card #230 (write-time integration-first default), card #236 (review-time necessity gate). This is the write-time REDUNDANCY filter neither covers. Target path given by the user: sports_inference/ubet-devenv/workspace/.claude/skills/orchestrated-feature-dev/ — but that is a CLI-synced consumer copy, byte-identical to AI-rules-repo/skills/claude-code/orchestrated-feature-dev/, so it will be overwritten on the next sync. Durable fix belongs in AI-rules-repo source + cursor/antigravity mirrors.
3 weeks ago
UBET-4355 — default chain migration to belief chainJira UBET-4355 "default chain migration to belief chain" (Task, Todo, Sprint 119, 4h, reporter Daniel Jiwoong Im, assignee Quan Vo, parent epic UBET-4282 "In-App Currency Betting"). NO description and NO comments in Jira — every requirement is derived from code. EPIC STATE at pickup: 4334 Done, 4335 Done, 4336 IN PROD, 4337 Review (PR #785 open), 4338 Review (PR #790 open, based on #785). Two NEW siblings created after 4338: UBET-4348 "Give 10 BP to new users once they register and complete category interest selection" (Todo, Quan) and UBET-4346 "in-app betting" (In Progress, Nick Mai — the FE half). So the epic is being driven toward launch. FINDINGS (read-only, from origin/main): 1. THE FRONTEND PICKS THE DEFAULT CHAIN BY LOWEST CHAIN ID. CommonServerDataProvider.tsx:72 does `Object.keys(chainsResponse.data.chains)[0]`, and chainRoute.ts returns `Object.fromEntries(...)` keyed on the chain id as a string. JS orders integer-like object keys ascending, so the first key is the numerically smallest chain id — not insertion order, and get_tokens.sql has no order by anyway. 2. THE BELIEF POINTS CHAIN ID IS INT4 MAX. beliefPointsTestHelpers.ts: BELIEF_POINTS_CHAIN_ID = 2147483647. 2147483647 is still a valid JS array index (< 2^32-1), so it sorts ASCENDING LAST behind 56/8453 (prod) and 97/84532 (testnet). The belief chain can therefore NEVER be the FE default as the code stands, even with beliefPointsConfig.enabled on. 3. THE LOGGED-IN PATH CANNOT REACH THE POINTS CHAIN AT ALL. All four resolution sites (HomePage:100, ActivitiesSidebar:22, RecommendedMarkets, utils.server.ts getInitialChainId:62) read `isAuthenticated ? wagmiChainId : (anonymousChainId ?? wagmiChainId)`, and wagmiChainId comes from SUPPORTED_CHAINS (constants.ts:50 = bsc/base or bscTestnet/baseSepolia), which can never contain the points chain. This is FE1 in tmp/UBET-4282/ticket-breakdown.md, never ticketed — likely UBET-4346's job. 4. THERE IS NO BACKFILL OF POINTS TWINS. insert_points_market_twin.sql is called from exactly two places in createMarketService.ts (:504 creation, :603 resurrection — and UBET-4338 D35 removes the resurrection one). Nothing ever twins a market that was already open when beliefPointsConfig went on. So on the day the belief chain becomes the default, the landing feed shows only questions created after the flag flipped. 5. Re-iding the chain is not a plain UPDATE. blockchain.chain_id is the PK (upredict_backend.sql:100) with `unique (name)`, and three tables FK to it with no `on update cascade`: collateral_token.blockchain_id (:109), the homepage blacklist (:293), and a query table (:810). A re-id has to be insert-new / repoint-children / delete-old under a temporary name. BLOCKED ON DATA: both supabase MCP servers (testnet and mainnet) fail every call with `TypeError: fetch failed`, so the live state is unverified — whether the BP chain is seeded per env, its collateral_token id, and how many open money markets lack a twin.
2 weeks ago
UBET-4339 add Discord + Twitch social login (committed on worktree branches; Supabase config + R26 decision outstanding)Ticket UBET-4339 "Adding Tiktok & Insta login" (Task, Todo, reporter Daniel Jiwoong Im, assignee Quan Vo). NO description and no comments in Jira. Read-only investigation, no code changed. HOW LOGIN WORKS TODAY: FE SignInDialog -> signInSupabase -> supabase.auth.signInWithOAuth (Supabase Cloud, project qbosayukigxebyotpnel, GoTrue v2.197.0) -> /auth/callback exchanges the PKCE code -> FE posts the Supabase JWT to BE POST /auth/socialLogin -> SocialAuthService verifies against the Supabase JWKS -> findOrCreateUserBySocialId keys the app user on the Supabase `sub` UUID. Separately AlchemyAutoLogin mints an Alchemy JWT (sub = Supabase user id) to provision the embedded wallet. BE NEEDS NO CHANGES. Every downstream identity is the Supabase `sub`, which is provider-agnostic. SocialLoginInfo.email is extracted but never read anywhere (verified by grep). Alias seeding uses user_metadata.user_name/name and already falls back to a random alias. FE IS SMALL: widen the provider union in features/auth/actions.ts (currently the literal "twitter" | "google" | "facebook"), add a button + svg + i18n key, optionally a PostHog flag mirroring FACEBOOK_LOGIN. TIKTOK — doable but needs a shim, not a config toggle. Not a Supabase built-in. Supabase Cloud does now support Custom OAuth2/OIDC providers (dashboard, unlimited on Pro), BUT TikTok uses `client_key` instead of `client_id` on BOTH the authorize and token endpoints, and Supabase reserves `client_id` as non-overridable. authorization_params can patch the authorize step; the token exchange is not configurable at all. So a plain custom-provider config cannot talk to TikTok — it needs a small OAuth-compliant facade we host that translates client_id -> client_key. TikTok also returns no email, so the provider needs email_optional:true, and AlchemyAutoLogin's `if (!supabaseUser?.email) return;` guard would silently deny those users an embedded wallet — that gate must move to supabaseUser?.id. No email also means Supabase cannot auto-link identities, so a user who previously signed in with Google gets a second account and a second wallet. INSTAGRAM — not buildable as specified. Instagram Basic Display (the only personal-account login path) reached EOL 2024-12-04. Its replacement, "Instagram API with Instagram Login", works only for Business/Creator accounts, and Meta does not approve apps using Instagram purely for authentication. Nearest available substitute is Facebook Login, which is already wired and enabled behind the `facebook-login` flag. ALSO: supabase-js must be bumped — installed @supabase/auth-js 2.87.1 has a closed Provider union with no `custom:${string}`; 2.116.0 adds it (and splits 'x' OAuth2 from the OAuth1.0a 'twitter' we currently use). TESTNET USAGE for context: google 153 identities, twitter 33, facebook 1.
last week

Blocked

0

Need Review

96
CCP environments: async build handoff — POST creates model + zipball only, add PUT for buildStatus/buildNotesVia orchestrated-feature-dev. Workspace: ccp-environments-async-build/tmp/environments-async-build/. Branch feat/environments-async-build off origin/main (worktree ccp-environments-async-build/). PROBLEM: POST /environments currently validates, creates the Environment model, clones the repo, audits the Dockerfile, AND builds the Docker image — all inside one web request. This does not work; the build must be out-of-band. TARGET WORKFLOW: user submits details -> CCP retrieves a shallow (HEAD-only) zipball of the chosen repo and stores it (currently a temp dir; should probably live in the data model) -> on its tick cycle a compute node finds a build request -> icp-shell-client pulls the zip, unzips, builds the container, and reports progress back through the ICP environments proxy to CCP. THIS TASK (CCP only): 1. POST /environments: validate inputs, retrieve zipball, create initial Environment model. Remove clone/audit/build from the request path. 2. Add validation: GitHub URLs must be HTTP(S) — SSH URLs unsupported. 3. Add PUT handler accepting buildStatus and buildNotes from the build process. 4. Update tests as needed. REFERENCE WORKTREES (read-only): - ccp PR #113 "Added zipRepo functionality" — MERGED, already in origin/main (utils/zipRepo.js) - icp-pr-37/ — ICP PR #37 "implemented environments proxy middleware" (middleware/environments.js) - icp-shell-client-pr-25/ — shell-client PR #25 "initial docker image build implementation" (modules/docker.py) Test command: npm run will_test (needs local Mongo :27017 + `docker compose up -d gcs-emulator`).
2 months ago
Quant: add Vietnam stocks (SSI FastConnect + VN30F1M) as a second marketUser (Vietnam resident) wants to extend quant-trading beyond crypto to Vietnamese equities, asking specifically about SSI's APIs. USER DECISIONS (2026-08-06, via AskUserQuestion): 1. GOAL = "Research now, trade later" — build the research layer but pick vendors assuming eventual execution. 2. INSTRUMENT = BOTH cash equities (cross-section) and VN30F1M futures. 3. ACCESS = has an SSI trading account (FastConnect key not yet confirmed); explicitly wants other vendor options kept open (vnstock/DNSE/TCBS), not an SSI lock-in. RESEARCH DONE THIS SESSION (needs writing up): - SSI FastConnect Data endpoints confirmed (DailyOhlc, IntradayOhlc 1-min, DailyStockPrice w/ ceiling/floor + foreign flow, Securities, IndexComponents, DailyIndex; pageSize max 1000; markets HOSE/HNX/UPCOM/DER/BOND). Auth = consumerID/secret -> AccessToken; RSA+SHA256 for trading. Requires SSI account + FIXED IP registration + 1-yr renewable term. Rate limits enforced but UNPUBLISHED. History depth UNDOCUMENTED — the #1 open question. - Free alternative: vnstock (VCI/TCBS/DNSE). DNSE caps minute data at 90 days, daily 10y. Legacy Vnstock class EOL 2026-08-31. - VN mechanics that break the engine: T+2 settlement hold-lock; no short selling in cash equities (SSC roadmap 2026-2028); ~60bps round trip (SSI iBoard 0.25%/side + 0.1% PIT on SELL proceeds regardless of P&L) vs crypto's 24bps; price bands +-7% HOSE / +-10% HNX / +-15% UPCoM; lot size 100; sessions 9:00-11:30 + 13:00-14:45 w/ ATO/ATC auctions; ~250 trading days/yr vs BARS_PER_YEAR["1d"]=365 (Sharpe inflated 1.21x). - VN30F1M: T+0, shortable, multiplier 100,000 VND/index point, IM 17% (VSDC), fees ~2,700 HNX + 2,550 VSD + 1,000-3,000 broker per contract, PIT = 0.1% x price x multiplier x qty x IM (~1.7bps effective). Circular 87 (eff. 2026-07-01) is the current derivative-tax rule. - TIMING FLAG: FTSE Russell upgrade Frontier -> Secondary Emerging effective 2026-09-21 — an index-inclusion flow event and regime break; must be embargoed in any backtest. STRATEGIC POSITION (from the repo's own track record): every single-asset directional arm has FAILED honestly (ML v0-v3, cross-section on 9 coins, funding-carry gate, pooled-panel). Only always-on carry ever beat B&H. So do NOT re-run single-asset direction on VN. The argued case for the port is CROSS-SECTIONAL BREADTH: ~400 liquid HOSE names vs 9 coins directly attacks the "~83-sample starvation" diagnosis, plus foreign-flow is a VN-specific factor with no crypto analogue. Engine transfers mostly unchanged (loop is index-based, assert_no_gaps is crypto-ingestion-only, every verdict runner already takes a bpy= override). What breaks: cost model is symmetric with no sell-tax, no settlement lock, no long-only constraint, no lot rounding, no ceiling/floor fill gate, BARS_PER_YEAR is 24/7.
quant-trading·main2 months ago
Mirror testnet reaction-enum migration into production for release 20260807.1 (upredict-infra #701)Production release PR https://github.com/SportsFI-UBet/upredict-infra/pull/701 upgrades production 20260805.2 -> 20260807.1. Today it ONLY bumps ci/configs/production/pipeline-config.json. Release contents (upredict-backend): #730 [UBET-4246] weekly ranking counts only markets created this week; #728 [UBET-4243] Subscribed notifications tab API. FINDING: production is missing Terraform/production/postgres/migrations/20260805000000_add_reaction_market_interaction_type.sql (present on testnet). It adds 'Reaction' to upredict_backend.market_interaction_type_enum and seeds notification_watermark for source_type 'MarketInteractionReaction'. Why it is required at 20260807.1 (not at 20260805.2, so prod is not broken today): - marketInteractionsWorker.ts inserts market_interaction_log rows with interactionType 'Reaction'; the column is market_interaction_type_enum, so the insert errors without the value. All of comment/reply/mention/vote notification inserts + push notifications run in ONE queryWithTransactionHandled, so the whole notification pipeline rolls back every tick. - appSql get_unified_notification_feed / get_unified_notification_unread_count / get_subscribed_notifications / update_market_interaction_log_read_subscribed all compare interaction_type against 'Reaction' -> invalid input value for enum, so the notification feed + unread count + the new Subscribed tab endpoint all 500. - The watermark seed matters too: without the row the worker starts at cursor 0 and would replay every historical comment_reaction. Why CI does not catch it: apply_sql.sh only wholesale-applies upredict_metric_views.sql, upredict_index_view_function.sql and config.sql. sql_schemas/upredict_backend.sql (which declares the enum) is never applied, and PostgresMigrationCheck/pgdiff is informational and does not diff enum labels. All 15 checks on PR 701 pass. NOT needed for this release: 20260806000000_add_follow_notification_log.sql and the followNotificationCooldownDays config key are only referenced at tag 20260807.2, which this PR does not deploy. Also verified in sync / no action: Terraform testnet vs production is structurally identical (last shared change 589a38a); config.ts unchanged across the release range so no new required config keys; production already has the point-ledger source types, user_follow, point_ledger.meta and user_alias migrations; frontend has not consumed the Subscribed notifications API yet. Precedent: PR #700 shipped the same shape (migration .sql + atlas.sum + config in the deploy PR). Sibling card #157.
2 months ago
claude-usage: repo drill-down is dead for claude-usage and quant-trading (label maps disagree)Reported 2026-08-10 with a screenshot: https://claude-usage.quanvo.dev/split/repo/personal%2Fquant-trading?preset=90d renders "No data in this range" even though quant-trading obviously has spend. # Measured on production (HTTP probes with the session cookie) Repo tab, single day 2026-08-08: - row `personal/quant-trading` = $136.00 / 1853 events - row `personal/claude-usage` = $124.54 / 1728 events - row `(unattributed)` = $9.22 / 111 events Drill-downs for the same day: - `/split/repo/personal%2Fclaude-usage` -> EMPTY - `/split/repo/personal%2Fquant-trading` -> EMPTY - `/split/repo/git-repos%2Fpersonal` -> $128.67 / 1783 events (a label the repo tab NEVER offers) - `/split/repo/(unattributed)` -> $3.09 / 45 events (its own row says $9.22 / 111) Per-day sweep: only 2026-08-02 and 2026-08-08 are broken; every window containing either of those days is broken, which is why 7d/30d/90d/all all fail and `today` works. Only claude-usage and quant-trading are affected — every other repo on those days drills down fine. # Root cause Two different code paths compute the repo label from two different row sets, and they disagree. - `costPerDay` (src/server/usage-queries.ts) groups by (day, projectSlug) and carries `repoKey: { $first: "$repoKey" }`. One slug can legitimately carry SEVERAL repoKeys (or a mix of present/absent) since per-turn attribution, so `$first` silently drops the rest — and it is unordered, so which one survives is arbitrary. - `repoKeyForLabel` (same file) groups over EVERY (projectSlug, repoKey) pair where repoKey exists. Both then apply `repoLabelsAcrossRange` (shortest projectSlug wins). Because the second sees strictly more pairs, it can pick a SHORTER label. Here the slug `git-repos/personal` (19 chars) beats `personal/claude-usage` (21) and `personal/quant-trading` (22), so the drill-down's map names both repos `git-repos/personal` while the dashboard still renders `personal/claude-usage` -> the label lookup finds nothing -> MATCHES_NOTHING -> "No data in this range". The `git-repos/personal` + repoKey pairing comes from `bin/sync.mjs:105-115`: the fallback attribution takes `projectSlug` from the project directory's own recorded cwd but `repoKey` from `resolveRepoKey(projectDir)`, so a turn with no evidence gets a parent-folder slug glued to a real repo's key. # Three defects, one cause 1. Dead link — the dashboard renders a drill-down link that resolves to nothing. 2. Under-counted repo rows — claude-usage's row is missing $4.13 / 55 events on 08-08; they fall into `(unattributed)` instead. 3. `(unattributed)` row ($9.22/111) disagrees with its own drill-down ($3.09/45) for the same reason. Latent: `$first` without a `$sort` is arbitrary, so repo labels can flip between runs. # Open decisions (blocking) - Naming rule when one repo has several folder names. Today "shortest wins" (D20), which lets a generic parent folder hijack a repo's name AND collide two repos onto one label. - Whether the sync CLI's fallback pairing is fixed too (stops new bad data; old rows only correct on a re-sync, since both fields are `$set`).
2 months ago
claude-usage: a container folder inside a repo steals its sub-repos' spend (ubet-devenv 40%)Operator report 2026-08-10 (screenshot, Repo tab, last 30 days): `sports_inference/ubet-devenv` is the #1 repository at $1971.32 / 12799 events. The operator does not work on devenv — it is a workspace container; the real work is in upredict-backend / -frontend / -infra and their ticket worktrees. # Layout `ubet-devenv` IS a git repo (SportsFI-UBet/ubet-devenv). Its .gitignore contains `workspace/*`. Inside `ubet-devenv/workspace/` live upredict-backend, upredict-frontend, upredict-infra plus ~45 ticket worktrees, each its own repo. All 63 UBet transcripts run with cwd = `.../ubet-devenv/workspace`. # Root cause Three pieces collide: 1. `resolveRepoAt` runs `git rev-parse --show-toplevel`, which walks up to the nearest .git and does NOT consult .gitignore. `workspace/` has no .git, so git answers `ubet-devenv` — correctly, for the question asked. 2. `attributeTurns` gives cwd precedence over files ("a turn run inside one repository while READING a file from another is working on the first"). 3. That precedence assumes a container folder resolves to NOTHING, letting the files decide. True for `~/Documents/git-repos/personal` (not a repo). False for `ubet-devenv/workspace`, which is inside one. Measured with the real parser over the real transcripts: 5,305 of 13,293 turns (39.9%) attribute to devenv. 5,014 turns record `.../workspace` as their literal cwd. Files cannot rescue it: of those 5,305 turns, 94.1% touched NO resolvable file (they ran a command or just replied). Only 4.2% would move on file evidence. # Rules measured (13,301 turns, devenv share) - current: 39.9% - A, cwd is gitignored by its repo: 2.1% — but depends on the operator having happened to gitignore the container; REJECTED by operator as too narrow (the personal dir has no .gitignore). - B, cwd directly holds nested repos: 16.4% — REJECTED, actively harmful. It discards a GOOD cwd on 3,297 turns in upredict-backend alone, because `upredict-backend/contracts` is itself a repo. Cannot distinguish "container of repos" from "repo that vendors a sub-repo". B+C (13.3%) scores worse than C alone. - C, ancestor never overrides: 4.7% — CHOSEN. # Chosen rule A repository that strictly contains another candidate never overrides it. If the cwd resolves to a repo whose root is a path-prefix of either (a) the repo the session is already in, or (b) the repo this turn's files point at, the cwd answer is discarded as less specific. Cheap (string prefix compare on two already-resolved paths), no filesystem scan, no .gitignore dependency, and it cannot discard evidence unless a MORE specific answer exists. Residual 4.7% is mostly session-start turns before a sub-repo is established, plus genuine devenv work. Left alone deliberately. # Scope Parser fix + spec + a full `bin/backfill.mjs` re-run to repair stored history (fields are last-writer-wins). The backfill does NOT need a deploy — it runs locally and re-uploads corrected attribution. Past days' figures will visibly move, and every project is re-tagged, not just UBet.
2 months ago
AI-Kanban: stop echoing the whole card on writes, add a lean resume view, cap entry lengthServer-side half of the post-compact context work (skill-text half was card #173). Three changes to the dispatch MCP layer, all in src/mcp/. Measured evidence from two real sessions: - Kanban tool results were 31.6% (claude-usage 33a644a2) and 43.1% (ubet-devenv d7543d8a) of ALL tool output — above Read and Agent in the second. - src/mcp/tools.ts toCardResult returns the ENTIRE card (full description + all decisions[] + all progress[]) from every write tool. Consecutive append_decision results on one card grew 597 -> 791 -> 965 -> 1332 -> 1625 -> 2934 -> 4905 -> 5314 -> 5847 tokens. The 9th decision cost 5,847t to record. - get_card_context on card #140 = 9,078t (12 decisions 4,479t + 8 progress 4,286t). Card #141 carried MORE decisions (15) in 2,907t — same schema, 3.5x better entry discipline. S1 - Writes acknowledge instead of echo. New toCardAckResult returning {id, number, status, decisions: count, progress: count}. Applies to append_decision, append_progress, update_card, set_status, claim_card, adopt_card, create_card, mark_decision_outdated. get_card_context keeps the full card. S2 - Lean resume view. get_card_context gains a view param defaulting to "resume": header + nextAction + ACTIVE decision headlines (no why) + latest progress note only + counts. Trims superseded decisions, all why prose, older progress notes, description, and the metadata tail. Estimated 9,078t -> ~2,400t on #140, 2,907t -> ~900t on #141. Add a detail fetch so an omitted why is one call away. S3 - Length validation on append_decision: decision <= ~200 chars, why <= ~400, ERR_VALIDATION reporting the actual length so the agent rewrites shorter rather than truncating silently. Pairs with the template added to pre-compact-flush in #173. Ordering: S1 and S3 are independent; S2 depends on nothing but touches the same file. Existing tests: src/mcp/dispatch-tools.test.ts, src/mcp/dispatch-server.test.ts, app/api/mcp/route.test.ts. CI = lint + tsc --noEmit + tests.
2 months ago
Card gate: revisit test + repo-scoped reuse-first lookup in ai-kanban-track-sessionStop the board over-creating cards in the ai-kanban-track-session skill. Two failures: (1) cards for work nobody revisits — #129 migration mirror, #130 and #148 code reviews, #150 one chart bug — because the gate is "substantive and multi-step", which is true of nearly everything an agent does; (2) one request fanned into several cards — #158/#159/#160 — because the only pre-create check is a keyword search for the SAME task, so nothing looks for neighbouring open work on the same repo. Design is two questions in order: WHETHER a card exists = "would you want to revisit this in a week?" (a card is a handoff device, default off); WHERE the work goes = list open cards carrying this repo's tag, reuse by default, split only when genuinely separate. Measured constraint: repo identity must come from git, NOT .ai-rules.json — claude-usage and quant-trading have no such file, and AI-rules-repo and personal-infra carry scope "personal" only, which identifies no repo. No AI-Kanban server changes needed since repo identity rides on tags. User constraint: NO enumerated skip list, the gate is a principle plus rationale. Plan in tmp/kanban-card-gate/PLAN.md, 8 steps; only step 5 (hook wording + tests/hooks/kanban-track.test.ts) has automated coverage. Split from #113 (need_review), which built the current search rung, forceNew narrowing and hook reminder for a different failure mode (cross-session duplicates of the same task).
2 months ago
FE guard: drop notification rows with an unrenderable MarketInteraction typePostHog issue 019f6aff-7918-7de2-9b3e-33a3e9105574 "Unknown MarketInteraction type". BE #728 (UBET-4243, 2026-08-07) started emitting interactionType "Reaction"; the shipped FE build had configs only for Comment/Vote/Reply/Mention, so MarketInteractionPageRow/MarketInteractionBellItem hit the !config branch, fired captureException on every render and rendered null. FE #685 (UBET-4166, merged 2026-08-12, main eb2342a) already added the Reaction enum member + config, so the specific row now renders. This card handles the GENERAL case the user asked for: no unrenderable interactionType should ever reach a renderer again. Two seams carry MarketInteraction rows and neither guards interactionType: - useUnifiedNotifications (All tab) -> filterRenderableNotifications guards only the OUTER NotificationFeedItemType (added by #701 / UBET-4057), not the inner interactionType. - useSubscribedNotifications (Subscribed tab) -> no filter at all, and every row there is a MarketInteraction (Reply/Mention/Reaction) - the highest-risk surface. Decision: guard only, do not add new rendering/copy (user's call). Scope: shared predicate in utils.ts used by both seams, plus the rawCount pagination pattern for the subscribed hook so a fully-filtered page does not end pagination early. Worktree: workspace/upredict-frontend-notif-guard, branch fix/notification-unknown-interaction-type, based on origin/main eb2342a.
upredict-frontend·fix/notification-unknown-interaction-type2 months ago
claude-usage: name a repo after its main checkout, resolved from git (worktree labels)Operator report 2026-08-12 (screenshot, Repo tab, last 7 days): the top row is `workspace/upredict-backend-ubet-4179` ($897.44) and another is `concrete_engine/ccp-environments-async-build` ($81.44) — both ticket-branch worktrees, named as if they were the repository. # Diagnosis (verified this session, read-only) The GROUPING is already correct. All 44 `upredict-backend-*` dirs under `sports_inference/ubet-devenv/workspace/` are real `git worktree` checkouts sharing one origin (`SportsFI-UBet/upredict-backend`), as is `ccp-environments-async-build` off `ConcreteEngine/ccp`. Same remote → same `repoKey` → `mergeRepoRows` already folds them into ONE row. That $897.44 is already the whole repository. What is wrong is the LABEL. `repoLabelsAcrossRange` (usage-queries.ts) names a repoKey group after the projectSlug with the MOST EVENTS in the window (card #166 D1), ties broken by shortest. A ticket worktree the operator lives in out-votes the main checkout and takes the repository's name. # Rejected: hard-coded name/prefix rules (the operator's opening ask) `workspace/upredict-backend-worktrees` sits in the same folder, shares the `upredict-backend-` prefix, and is a DIFFERENT repository (`SportsFI-UBet/ubet-devenv`). A prefix rule would swallow it. This is the exact false positive D20 rejected name-matching for, present in the real corpus. # Chosen approach — ask git for the canonical name `git rev-parse --git-common-dir` from a worktree returns the MAIN checkout's `.git` (verified: resolves to `…/workspace/upredict-backend/.git` and `…/concrete_engine/ccp/.git`); a main checkout answers the relative `.git`. Resolve the main checkout root at sync time, store its two-segment slug as a new `repoName` field alongside `repoKey`, and have the label prefer it. No name heuristics, no list to maintain, works for every repo on every machine. Write path touched: `src/parser/repo.mjs` → `attribution.mjs` → `events.mjs` → `api/sync/route.ts` allowlist → `usage-store.ts`; read path: `usage-queries.ts` label ranking. # Known open items - Backfill re-run vs forward-only (D4 precedent on card #166: full `bin/backfill.mjs`, 2,532 files / 75,572 events). - Bare-repo-plus-worktrees layout has no main checkout to name. - Three stale comments + one stale test title still claim the label is the SHORTEST slug (usage-queries.ts:963/1184/1267, usage-queries.test.ts:155) — card #166 changed it to most-events and left these behind. Follows card #166 (label rule) and #169 (container-folder attribution).
claude-usage·main2 months ago
Fix orchestrated-feature-dev plan-format drift (stale rule + stale skill + competing templates)Diagnosis: since 2026-07-12 the orchestrated-feature-dev Phase-2 plan node stopped producing the create-implementation-plan format (## Technical Design + ## Behaviors to Implement + test checkboxes) and instead emits an AC/Test-Type/Depends-on format. Measured across 49 tmp/*/implementation-plan.md files: 100% compliant before 2026-07-12, progressively non-compliant after. Four causes: 1. .claude/rules/feature-development-guide.md was DELETED from ai-rules source on 2026-07-08 (commit 9b94992) but installed copies were never pruned — 15 copies still on disk, incl. the workspace root, so it is auto-loaded as project instructions into every session AND every sub-agent. It prescribes the exact deviating format and its "What NOT to include" list forbids the skill's Technical Design section. 2. feature-development-workflow was renamed to feature-dev-lite on 2026-07-12 (commit 4997720); 11 stale skill dirs remain installed, carrying the same competing template. Rename date == drift start date. 3. feature-dev-lite/SKILL.md itself restates the competing AC/Test-Type template (a genuine design collision, not staleness). 4. node-plan.md says "Use @create-implementation-plan" (prose, not an invocation) and never overrides that skill's Step 0 "ask the user for an identifier" + Step 1 "MUST stop and wait for the user" — un-followable for a sub-agent. User decision: do NOT add prune-on-sync to the CLI (prune risks deleting skills they still want).
2 months ago
RISE-15577 — Add SM org type + status columns to organizations_associations and backfillGive RSC a place to store Supplier Management's own organization type and status, per ecosystem, and backfill what SM already knows. SCOPE (operator-confirmed, narrower than it looks): migration + one-time backfill script ONLY. No write path, no GraphQL, no UI. RISE-15509 (card #171) still owns wiring POST /data-management/add-org-into-ecosystem to populate sm_type; the consumer/filter tickets (15478/15479/15510/15557/15487, all BACKLOG) own reading it. Not demonstrable alone — a prerequisite, like 15509. SHAPE: two new columns on organizations_associations (the `initiator` row = RSC's per-ecosystem org record), both new PG enums in SM's own vocabulary — sm_type = B R S V F I O, sm_status = draft/awaiting approval/active/inactive/suspended/blacklisted/closed. Matches SM_ORGANIZATION_TYPES / SM_ORGANIZATION_STATUSES and the SmOrganizationType / SmOrganizationStatus SDL enums RISE-15476 already shipped, so no translation layer. BACKFILL: reads SM's true values from cdc_passport_be.ecosystem_organization via ecosystem_organization_sync (product='rise'). Covers 11,192 of 285,703 initiator rows (3.9%); the rest stay NULL. Deliberately NOT derived from RSC's own columns — organizations.type is a lossy collapse (839 SM Brands stored as `retailer`, 9,184 Suppliers as `partner`, 1,084 Vendors as `external`) and organizations_associations.status is `active` on 285,277 of 285,284 rows, so a literal prefill would write known-wrong values and zero information. MR needs the [MIGRATION] tag. Design context: RISE-15509_SM_RSC_ORG_MODEL.md (repo root) and rs-backend-RISE-15509/DECISIONS.md D23–D32.
rs-backend·feature/RISE-15577/add-sm-org-type-status-columns2 months ago
Replace syntax-recipe exercises in Data Structures &amp; Strings (Python + C++)Follow-up to card #97. Operator's complaint: several exercises in the Data Structures &amp; Strings lessons are implementation recipes, not problems. Operator's test (corrected from my first pass): an exercise fails when the PROBLEM STATEMENT prescribes the implementation steps AND that machinery is not needed to produce the stated output. "Print it backwards" is a goal → fine. "Store in a tuple, print the tuple, unpack it, print each part" is a recipe, and print(f"({x}, {y})") gives the same output → fails. HARD FAILS (machinery provably redundant) — replace with new problems: - Python pool: #1 Swap Two Numbers, #2 Coordinates with a Tuple, #13 Build a Date with join, #15 Total and Average Score - C++ pool: #2 Sum of a Vector, #12 Split into Words (echoes input), #14 Total and Average Score, #18 Convert Between Text and Numbers, #20 Receipt Line - Lesson pages: Python inline Ex1 Swap, C++ inline Ex7 Split into Words SOFT FAILS (real goal, but statement names the tool) — reword to state only the goal: - Python pool #9, #16, #18, #20; C++ pool #3, #10, #15; Python inline Ex8; C++ inline Ex2, Ex6, Ex8 Constraint chosen by operator: KEEP feature coverage. Replacements must still genuinely require unpacking / join / stoi / to_string / vector / substr etc. — but require them for real, not as ceremony. Files: - app/lesson/programming-python/data-structures-exercises/exercises-data.ts (27 exercises) - app/lesson/programming-cpp/data-structures-exercises/exercises-data.ts (27 exercises) - app/lesson/programming-python/data-structures-strings/page.tsx (inline practice, L306-441) - app/lesson/programming-cpp/data-structures-strings/page.tsx (inline practice, L279-422)
2 months ago
RISE-15478 — Request Single Assessment: filter Assessed Org / Selected Partner / auto-share by standard's org type & statusOrchestrated-feature-dev run for RISE-15478 (epic RISE-15475, blocked-by RISE-15476 which is already in pre-prod testing; FE part of 15476 merged to master). Make the SINGLE assessment request form honor the standard's per-stakeholder organization type/status filters that RISE-15476 stored: - AC1 Assessed Organization dropdown filtered (All + Activation List tabs), with guidance text when a search matches nothing. - AC2 Selected Partner auto-display + "Add more" dropdown both filtered. - AC3 Assessment Visibility / auto-share shows "Limited to Organization Types: {types}"; Open-audit variant filters its dropdown too. - AC4 validate on save + revalidate on edit (New status only) → "Invalid organization" per stakeholder + removal alert. - AC5 follow-up assessment copies stakeholders then validates the same way. Gated by DF `release_org_type_sm` (precondition `supplier_management`) — must be fully inert when off. OUT OF SCOPE: bulk request (RISE-15510, backlog) and the Standard Settings UI itself (RISE-15476, shipped). Worktrees: - BE /Users/quan.vo/Documents/git-repos/inspectorio/rs-backend-RISE-15478 on feature/RISE-15478/request-asm-org-type-status-filter, based on the UNMERGED feature/RISE-15577/add-sm-org-type-status-columns (carries the 15476 BE commits + the SM org type/status columns on organizations_associations). - FE /Users/quan.vo/Documents/git-repos/inspectorio/rs-frontend-RISE-15478 on the same branch name, based on origin/master @ cab72cac2f. Workspace: rs-backend-RISE-15478/tmp/RISE-15478/ (TASK_CONTEXT.md holds the full ACs and the 17 QA test cases).
2 months ago
UBET-4296 tagging search broken — @mention search can't find 89% of usersTicket UBET-4296 "Tagging search broken" (Task, Todo, reporter Daniel Jiwoong Im, assignee Quan Vo). NO description and no comments in Jira. Repro from user: the @mention typeahead cannot find user "quan vo". ROOT CAUSE (confirmed on testnet, 2026-08-21). get_users_search.sql filters on exactly two fields: `au.nickname` and social `raw_user_meta_data ->> 'user_name'`. It never touches `au.user_alias`, and never touches the social `name`/`full_name` that Google populates. `user_name` is a Twitter/X-only field. Testnet counts (abstract_user LEFT JOIN auth.users): twitter 32/32 searchable; google 2/145; no-social (wallet/email) 8/219. Total 42/396 searchable — 354 users (89%) can never be returned by the mention search no matter what is typed. Meanwhile `user_alias` is non-null for all 396. All three "quan vo" rows (ids 27387/63158/102412) are google: nickname NULL, social user_name NULL, user_alias quanvo/QuanVo2/QUANVO3, social name "quan vo"/"Quan Vo"/"QUAN VO". Replaying the production query verbatim for q='quan' returns [] — reproduced exactly. SECOND, SAME-ORIGIN BUG: usersSearch.ts maps `username: row.rawUserMetadata?.user_name ?? null`, hand-rolled instead of using the shared convertRawMetadata() in socialUserInfo.ts, which does `user_name ?? name`. The mention route is the one place in the codebase that drops the `name` fallback. So even once search matches, FE getUserDisplayName({nickname, username}) gets two nulls for a Google user and renders "Unknown". THIRD (independent, FE): getMentionQuery in src/features/comments/utils.ts uses /(?:^|\s)@([\p{L}\p{N}_.-]*)$/u — no space in the class. Typing "@quan " closes the dropdown, so no two-word name is typeable in full. FOURTH (aggravator): `order by au.nickname asc nulls last limit 20` deprioritizes exactly the null-nickname users who dominate the table. Fix-variant counts measured on testnet for q='quan' / q='quan vo': current 0/0; alias-only 5/0; alias + social name 5/3.
last month
SEO image for reel markets captures TikTok cookie-consent dialog instead of the videoDiagnosis only, no code changed. SYMPTOM: the generated SEO/OG image for a TikTok reel market shows TikTok's "Allow cookies from TikTok on this browser?" consent modal where the video should be, over the two outcome buttons. CHAIN: - upredict-backend/typescript/services/seo_image_generator/src/index.ts — Lambda (cron every 5 min, Terraform/environments/backend/main.tf) launches puppeteer headless Chrome with userDataDir=/tmp/puppeteer_<uuid>, goes to {originUrl}/markets/{id}/place?preview=true, waitUntil networkidle0, waits for .preview-content, screenshots that ELEMENT as jpeg, PUTs to S3, inserts upredict_backend.market_seo_image. - upredict-frontend origin/main: PlaceBetPage.tsx:31-47 reads ?preview and adds .bet-card-preview; PlaceBetCardContent.tsx:404 is the .preview-content wrapper -> Field component=MarketOutcomes -> Market2TextBasedOutcomes.tsx:114 <MarketMainImage> -> MarketMainImage.tsx:15-21 returns <MarketReelEmbed> when market.reel is set -> MarketReelEmbed.tsx renders <iframe src=https://www.tiktok.com/player/v1/{videoId}?controls=1&autoplay=1> (utils.ts:252). - So the third-party TikTok iframe sits INSIDE the screenshotted element. ROOT CAUSE: a fresh userDataDir per invocation means every run is a first-time visitor to tiktok.com with no consent cookie, so TikTok's player embed renders its consent gate instead of the video. Nothing in preview mode suppresses the embed — .bet-card-preview (globals.css:314-322) only hides .market-outcome-participants, .market-outcome-footer and .market-progress-clock. networkidle0 + .preview-content are both satisfied while the consent modal is up, so the screenshot is taken and stored as a success. SECOND-ORDER: the backlog query selects only markets with NO market_seo_image row, so a bad image is written once and never regenerated. RELATED TICKETS (epic UBET-4107 Reel Market): UBET-4123 "download thumbnail for reel markets" (Todo) is the natural fix. UBET-4047 spike (Notion Embedded Reels Spike). UBET-4279 homefeed (In Progress), UBET-4280 create-market (Committed), UBET-4281 profile (Todo). SEO pipeline tickets: UBET-3650 seo render in backend, UBET-3818 Chrome in devenv, UBET-3613 (earlier same-shaped bug: SEO snapshot missed images).
last month
UBET-4323 render reel thumbnails in SEO images + Instagram via Meta tokenBug UBET-4323 (parent epic UBET-4107 Reel Market, Sprint 117, assignee Quan Vo). Follows the diagnosis on card #212 and the probe work on card #208 (UBET-4124). PROBLEM: SEO/OG images for reel markets screenshot the live market page under ?preview=true with a fresh puppeteer profile each run, so the embedded provider iframe renders its consent/login gate instead of the video. FIX: under ?preview=true render a static thumbnail instead of MarketReelEmbed, removing third-party iframes from the image pipeline. YouTube (URL-derivable) and TikTok (public oEmbed thumbnail_url) need no token — 47 of 77 testnet reel markets. Instagram (29 of 77) needs a Meta app token for Graph instagram_oembed (oEmbed Read review). The same Meta token also unblocks Instagram in the reel liveness probe, where every instagram.com reel is currently retired as terminal `unprobeable` on sight. Those retired markets must be re-armed or they stay dark after the token lands. ALSO: bad images are never retried — the backlog query skips markets that already have a market_seo_image row (75 of 77 on testnet). Decide whether to delete those rows, and whether that applies to production. Following orchestrated-feature-dev. Workspace tmp/UBET-4323. Worktrees: upredict-backend-ubet-4323 and upredict-frontend-ubet-4323, both on branch feat/UBET-4323-reel-seo-thumbnail (BE stacked on feat/UBET-4124-reel-play-error-notification at 2aa1d186; FE on origin/main at 67c81b7).
4 weeks ago
Code review: upredict-backend PR #776 — UBET-4317 associate existing markets into category listReview https://github.com/SportsFI-UBet/upredict-backend/pull/776 ([UBET-4317] Associate existing markets into category list, author quangtran-sportsinference, branch UBET-4317-associate-existing-markets-into-category-list -> main). 47 files, +6150/-0. WHAT IT DOES: adds a new `market_category_classification` Python service that classifies markets against an interest taxonomy via a batched LLM call (bisects a failing batch to isolate one bad market; bounded by a Lambda-derived wall-clock deadline). Persists per-market classification vectors as safetensors matrices in S3 with positional `row_index` identity in Postgres, precomputes per-category/subcategory relevance scores, and adds a belief_locker TypeScript consumer (get_sampled_markets_for_categories.sql, get_user_picked_category_ids.sql, marketRecommendationService.ts) for home-page category recommendations. 7 new tables. SETUP: worktree at workspace/upredict-backend-pr-776; review artifacts under its tmp/pr-776/review-changes/ (HOLISTIC.md, DIFF.patch, LENS_*.md), final report at tmp/pr-776/review-changes.md. Running /review-changes at fan-out depth — all six lenses applicable, none skipped. HOLISTIC's headline concerns for the lenses to resolve: 1. `row_index` is positional identity split non-atomically across S3 and Postgres; write_classification_rows' `on conflict do nothing` can produce a non-contiguous orphan the tail-truncation repair mis-handles. 2. `_write_relevance` deletes and re-inserts the entire markets x categories cross-product every run on an explicitly unmeasured cost assumption. 3. The vertical slice is incomplete in a way the PR description overstates: sampleMarketsForUserPicks has no caller, user_interest_category is read but written by nothing here, and the taxonomy rows the dense-dimension_index assertion depends on are seeded outside this repo. 4. Nothing ever reclassifies a market (prompt_version written, never read). 5. The new sampling query lacks the deadline/status filter get_home_markets.sql goes to lengths to get right. Also noted as done WELL (don't re-litigate): bisection instead of dropping whole batches, real Lambda-derived deadline, refusing to run on an empty taxonomy, stabilized softmax, make_conninfo over f-string conninfo.
4 weeks ago
UBET-4334 — Belief Points as a chain, a token and a balanceOrchestrated-feature-dev run for UBET-4334 (parent epic UBET-4282 "In-App Currency Betting"; blocks UBET-4335 "Every market also exists in belief chain"). Sprint 118, assignee Quan Vo, 4h estimate. TICKET ASKS (4): 1. Add a `blockchain` row for Belief Points + one `collateral_token` under it. 2. Tell the two token kinds apart with a flag; make `collateral_token.address` optional. 3. Let a `market` row exist with no `contract_address` and no `deadline_block`. 4. Expose a user's spendable balance = frozen legacy snapshot + sum of points history, computed on demand. AC: /chain returns the new chain and its token; a market row can exist against that token with no chain fields set; a balance reads back as the same total the profile already shows. PRE-READ FINDINGS (main session, before orchestration): - Ask 4 likely ALREADY EXISTS: GET /userRealtimePoints -> get_user_realtime_points.sql already computes greatest(0, legacy_leaderboard_view.legacyPoints + sum(point_ledger.points) post-cutoff) on demand. Open question whether "spendable" means net of points locked in open bets. - AC "/chain returns the new chain" will NOT happen from DB rows alone: chainRoute.ts:71-80 silently drops any chain absent from the predictionMarketAddresses map, which is built in index.ts:88 collectPerChain via createEvmService -> asserts an Alchemy URL exists and chainId is in `chains` (evmService.ts:105-107). A synthetic points chain crashes the service at boot. Requires a chainRoute code change + a decision not written in the ticket. - Nullable collateral_token.address breaks /chain for EVERY chain: get_tokens.sql selects ct.address, parsed with z.string() at chainRoute.ts:59; one NULL throws and 500s the endpoint. - ChainResponse.tokens[].contractAddress is `string` and marketContractAddress is `EthAddress`, both non-nullable (jsonSerializable.ts:16,20) — widening is a frontend contract change + codegen regen. - `unique (blockchain_id, address)` stops enforcing once address is nullable (Postgres NULLs are distinct) — needs a partial unique index or NULLS NOT DISTINCT. - Nullable market.contract_address / deadline_block touches ~13 and ~6 appSql files; metric views select+group by contract_address so points markets become a NULL bucket rather than dropping rows (degrades, not breaks). - 4h estimate looks light. RELATED PRIOR WORK: card #211 (6a993f86c0af1b6c8a13d4b4) UBET-4282 currency model design brainstorm (staled) — read for prior decisions.
4 weeks ago
UBET-4335 — Every market also exists in belief chain (points twin)Orchestrated-feature-dev run for UBET-4335 (epic UBET-4282 "In-App Currency Betting"; blocked by UBET-4334, blocks UBET-4336). Sprint 118, assignee Quan Vo, 3d estimate. Ticket ask: a market created on a real chain also gets a second `market` row on the Belief-Points collateral against the SAME `market_spec` — question, outcomes, images, translations and creator shared, one extra row. Covers already-open markets (backfill) and resurrection. Open decision in the ticket: creation-transaction hook vs worker. Worktree `workspace/upredict-backend-ubet-4335`, branch `feat/UBET-4335-points-market-twin`, based on `feat/UBET-4334-belief-points-chain` @ 5bf32288 (NOT main — 4334 is unmerged and this depends on `blockchain.kind`). Pre-read ticket check at `workspace/tmp/UBET-4335/TICKET_CHECK.md`. Headlines from it: - Ticket text is STALE on three bullets: 4334 chose sentinels not nullable columns (D3), so `deadline_block` stays `not null` with the sentinel 2147483647 (D14) — "no block number" is wrong; the `kind` flag is on `blockchain` not `collateral_token` (D6), reached via market -> collateral_token -> blockchain. - MISSING from the ticket, biggest item: `get_search_markets.sql:81` dedups `distinct on (m.spec_id)` and ties break on `m.id desc`, so the later-written twin ALWAYS wins. Search has no chain filter at all — every twinned question would return its points market to every user on every chain. - MISSING: creator page + home feed both gate on `chainId is null or ...` and both routes pass `query.chainId || null`, so with no chain filter a question lists twice. - MISSING: `resurrectMarkets()` loops `this.perChain.keys()` (createMarketService.ts:521); `perChain` entries need an `EvmService`, so the points chain can never be in it. Resurrection needs its own path, and `createMarket()` cannot be reused for the twin (it starts with `getLatestBlockInfo`) — must be a direct insert. - 4334 deferrals R9 + R10 DISSOLVE: `market.contract_address` is chain-wide (`chainData.predictionMarketAddress`), so the blacklist and reward aggregates already pool per contract for real chains. Still live: D20 (`BigInt(token.maxBetSizeWei)` FE crash, made reachable by this ticket) and D22 (points market vanishes from listings after deadline until BE4 settles it). - Ticket's two stated warnings are real but described off: the auto-vote is `on conflict (user_id, market_spec_id) do update` (insert_bet.sql:45) so the second bet OVERWRITES the first vote; the weekly ranking pools bet COUNTS not volume (upredict_index_view_function.sql:1053). - Recommendation on the open decision: worker (only option that covers resurrection at all, covers the backfill with the same code, idempotency is a query not a lock, precedent in tasks/copyMarkets.ts + ResurrectMarketsTask). Note CopyMarketsTask is NOT reusable — it calls createMarket() and duplicates the question.
upredict-backend·feat/UBET-4335-points-market-twin4 weeks ago
Test review: add a necessity check, name the mock-entailment signature, cut the assert-on-mocks guidanceReview-time counterpart to card #230 (which set the WRITE-time integration-first default). Trigger: the user keeps finding unit tests that guarantee nothing — they mock an API, then assert the API returned the mocked value. The assertion is entailed by the test's own arrange block, so production code is not in the causal path. Audit finding: every test-review section asks "is this test MEANINGFUL?" (answer = improve the assertion) and none asks "should this test EXIST?" (answer = delete it). Worse, test-quality-reviewer actively manufactures the bad pattern — its Sensitivity pillar recommends `expect(fn).toHaveBeenCalledWith(...)` and its Resilience pillar says "checking internal helper calls is OK for unit tests". review-changes and feature-dev-lite both defer their deep test pass to that skill. Scope (all 3 agent variants — hand-maintained ports, no generator): 1. node-lens-tests.md — necessity check FIRST, ahead of coverage; name the entailment signature with a concrete bad example; SHOULD FIX severity floor. 2. lens-common.md — "Suggested fix" must admit deletion; add a tests-lens accepted failure-mode form (the MISSED DEFECT that ships green), or a useless test gets filed as a NIT maintainability note. 3. orchestrated node-validation.md §2 — same necessity check + entailment signature. 4. test-quality-reviewer — CUT the mock-interaction sensitivity lines and the "internal helper calls OK" clause outright (user: remove the lines, do not add a caveat). Especially wrong for integration tests. 5. feature-dev-lite — add the inverse to "What to Avoid" so it is prevented at write time too. Do NOT hand-edit the .claude/skills/ dogfood copies — the CLI regenerates them.
3 weeks ago
Plan + build Files, Sorting & Records handouts (Python + C++) — the 100-student problemNext lesson after Data Structures & Strings (cards #97, #191). Covers curriculum module 8 (Files) plus sorting and records, but organised problem-first, not topic-first. # Method (operator's, corrected through several passes) One real contest problem drives everything: read 100 students, print lowest→highest score. Every technique appears ONLY when a wall demands it. Second thing taught is decomposition (GET/KEEP/DO/SHOW) — NOT to be called "divide and conquer" (that name is reserved for mergesort/binary search later). Walls, in order of pain: 1. Can't retype 100 rows → file input (ifstream / open) 2. Scores sort fine, but the problem wants NAMES → parallel arrays produce a program that compiles, runs, looks right, and pairs everyone wrong (shown on 5 rows, never 100) → tuple/pair in a list/vector 3. Payoff: box-by-box diff shows GET/KEEP/SHOW changed, DO is byte-identical. Verified by compiling both C++ versions. # Corrections the operator made during design (do not regress) - Sorting must come BEFORE records. Without sorting there is no wall — parallel arrays work fine. - The name requirement is ADDED to a solved problem (requirements grow), not "restored after simplifying". - No session labels in the handout — one continuous flow, operator splits it live. Break points after sections 5, 6, 10. - Do NOT teach shell redirection (`./prog < scores.txt`). I invented a fake "spoiler" conflict around it across two turns; it is not in the syllabus and adds a failure mode for nothing. - Section 0 = install only (bash, language toolchain, compile+run). No proof checklist, no troubleshooting section. - Dropped the "four-box worksheet" as a graded artifact — the boxes stay as explanation only. # Decisions - Windows only. No Mac/Linux track. - TWO separate files, one per programming language (both still bilingual EN/VN per repo convention). - Python setup: Git Bash + python.org installer ("Add python.exe to PATH"). - C++ setup: MSYS2 ONLY — its UCRT64 shell is itself bash AND has g++ on path, so students never edit Windows PATH (the classroom-killer step). `pacman -S --needed base-devel mingw-w64-ucrt-x86_64-toolchain`. Verified against current VS Code / mingw-w64 docs. - Data file uses ONE-WORD names; full names with spaces break `>>` and would cost half the lesson on stringstream. Save as a later bump. - Deliberately NOT taught here: comparator lambdas / `key=`, `struct`, dicts-in-a-list. Field-order choice + reverse covers everything. Comparators arrive when a mixed-direction sort (score desc, name asc) demands them. # Done so far - app/lesson/programming-python/files-sorting-records/manifest.md - app/lesson/programming-cpp/files-sorting-records/manifest.md # Remaining - page.tsx for each (print-friendly component structure, CodeBlock language py/cpp) - scores.txt generator (100 Vietnamese names + random scores, regenerable for homework variants) - Add both to courses.md and app/page.tsx nav - One clean Windows run-through of section 0 before it goes to students (I am on macOS and could not test Git Bash / MSYS2) # Stint 2026-09-27 — more section 9 exercises Operator asked for more exercises in section 9 and more sorting of records (struct / tuple). Adding: 3 sort-order exercises on name+score (highest first, top 10 with place, closest to 75) and a three-field part (name math literature): Python 3-part tuple, C++ struct taught briefly in section 9, with sort-by-math / sort-by-literature / sort-by-total exercises. Homework (alphabetical) stays last.
3 weeks ago
UBET-4338 — Re-check every market and bet surface for points (BE5 sweep)Assessment only, nothing implemented — no branch, no worktree. Reading pass over Jira UBET-4338, its four sibling tickets, and the board records from implementing them (cards #211, #226, #228, #240, #246). TICKET. UBET-4338 "Re-check every market and bet surface for points", Task, Todo, Sprint 118, assignee Quan Vo, 1d4h estimate, parent epic UBET-4282 "In-App Currency Betting", blocked by UBET-4337. It is BE5 in tmp/UBET-4282/ticket-breakdown.md. Not an open audit: the PO asked "do we need a BE5 to sweep everywhere a market is shown" on 2026-09-08 and the answer was to make it a DEFINED CHECKLIST sequenced after BE4. The seed checklist already exists in tmp/UBET-4282/JOURNAL.md under "Survey — what breaks on market and bet read surfaces". EPIC STATE. 4270 planning Done. 4334 Done (PR #779 merged 2026-09-11). 4335 Done (PR #781, on main at a924580b). 4336 Jira Review, PR #783 OPEN — not merged. 4337 Jira In Progress, branch feat/UBET-4337-belief-points-settlement unpushed, card #246 need_review. So two of 4338's four prerequisites are not in main; the sweep has nothing complete to walk. CARRIED FORWARD INTO THIS TICKET (from the board): 1. Creator "fees collected" — card #246 D20 (R53) left explicitly to BE5, no code and no tests. Surface is get_creator_and_total_collected_fee. 2. Homepage blacklist — JOURNAL survey: not exists on lower(contractAddress) passes a points market, and market_homepage_contract_blacklist is keyed on contract address, so a points market cannot be blacklisted from the homepage. Compounded by card #226 D14/D23 (R10): the blacklist PK is (contract_address, blockchain_id) and every points market shares the 0x0 sentinel, so blacklisting ONE points market hides ALL of them. Card #226 D23 says this goes live the moment any real points-market row exists — 4335 is merged, so the infra token seed opens it. Never resolved by 4335/4336/4337. Possibly overlapping with 4335 D18 (R32 suppression list made an explicit no-op) — needs checking whether they are the same list. 3. get_referrer_summary silently returns zeros for points markets (left join on lower(m.contract_address) never matches the sentinel). JOURNAL calls it "probably correct, but it should be a deliberate answer, not an accident". No decision recorded anywhere since. 4. Resurrection first activation — card #228 D17 (R18): the resurrection twin hook is PRODUCTION-DORMANT because nothing can move a points market off Open until 4337 lands. 4337 gives points markets a close, so that path fires in production for the first time inside 4338's window. 5. Deadline disappearance — card #226 D22: a points market silently vanishes from home and search once its deadline passes (get_home_markets open/closed clause unsatisfiable), root cause "no settlement path". Deferred to 4335, still unverified. 4337 supplies the settlement path; the sweep should confirm it actually fixed it rather than assume. 6. Stranded-bet sweep — chain of deferrals ends unconfirmed: card #240 D52 dropped the EVM-only filter on update_orphaned_bets ("belongs to 4337"), card #246 D27 then decided no points guard is needed because atomic settle leaves nothing stranded, marked "pending user confirm". 7. Collateral-token id collision — card #226 D21: the check that the seeded BP token id does not collide with copyMarketsConfig.sourceCollateralId/targetCollateralId was never performed. Belongs on an infra checklist, not code. FRONTEND — the whole other half, and no tickets exist for it. Card #211's nextAction still reads "create the 3 FE tickets if they should be tracked separately"; FE1-FE4 are drafted only in ticket-breakdown.md. Recorded FE breakage: - Card #226 D4: an unknown chain id throws at render in ChainDropdownItem/AnonymousChainButton with no error boundary, breaking chain switching for ALL users on ALL chains. Card #226's nextAction says flipping beliefPointsConfig.enabled on testnet is exactly what makes this real. - Card #226 D20: Token.maxBetSizeWei is typed non-nullable and the FE does BigInt(token.maxBetSizeWei); 4334 emits null for BP, and BigInt(null) throws (probe-confirmed) in MarketOutcomeAmountInput. A second, distinct crash site. - CHAIN_ICONS has no fallback. - The chain dropdown does not work when logged in at all: ChainDropdownItem calls setAnonymousChainId and never wagmi switchChain, while four resolution sites (HomePage, RecommendedMarkets, ActivitiesSidebar, server getInitialChainId) read `isAuthenticated ? wagmiChainId : anonymousChainId`. A logged-in user can never reach the points chain. - Card #228 D6: chain-aware search is backend-only; features/search/fetches.ts still sends no chainId. Deferred to an FE ticket that was never created. ALREADY DECIDED — the sweep must NOT re-raise these: - Points bets excluded from every leaderboard bet stat, wins/losses/volume/profit (card #240 D25), and from /activitySummary volume (D12). - Stake source type left out of every per-bucket leaderboard figure, so buckets stop summing to the headline while a stake is open (card #240 D8). - Betting both twins of a spec overwrites the public vote with the latest bet's outcome (card #240 D10, regression-tested). - Creator "my markets" shows a prediction twin as Pending because the winner lives on the money market (card #246 D46, user call, out of scope). - Weekly board cross-week effect accepted (card #246 D20, R50). - The chain is a hard partition: a user on BNB sees neither points markets nor their own points bets in history. Accepted deliberately. - Amount formatting needs no work: the FE uses viem formatUnits(value, token.decimals) in 7 files and the BP token is decimals=6. - Belief-points markets DO get their own generated SEO image (card #228 D40 dropped the exclusion) — a render surface worth walking, not a gap. WATCH: card #228 D20 made the weekly market ranking money-markets-only, but card #240 D51 later removed the kind='Evm' filter in generate_market_spec_points so points bets DO raise a question's ranking score. Check those two still mean what was intended together.
3 weeks ago
Quant: test the copy-trade TP-ladder geometry (+0.5/1.0/1.5% scale-out, -0.8% stop, 4h cap)Operator copy-trades a leader on Binance perps at ~100x: TP at +50%/+100%/+150% of margin, SL at -80% of margin, fees ~10% of margin per round trip. Translated to PRICE space that is TP +0.5%/+1.0%/+1.5%, SL -0.8%, round-trip cost ~10-12 bps — and the fee figure independently confirms ~100x taker fills. WHY THIS IS NOT ALREADY REFUTED: the live paper accounts 4/5/6 are `mean_reversion` lookback 60/240, entry_z=1.5, exit_z=0.0, take_profit_pct=0.0 — they exit on reversion to the mean, a target of roughly 15 bps, below the ~19-24 bps cost line that docs/research/ml-v1-improvements/04-frequency-and-fees.md says you must clear. The leader's targets are 50-150 bps, 3-10x larger. Accounts 4/5 losing ~13 bps/trade does NOT refute the leader's geometry; it confirms the cost doc. MEASURED (2026-09-21, scratchpad, real 1m bars 2020-01..2026-05, entry every 60m, hold<=240 bars, SL-wins-intrabar): - ZERO-SKILL entry: gross +/-0.1 to 0.3 bps per trade on BTC/SOL/DOGE, i.e. the ladder is a martingale. Net = -12 bps all-taker, -8 bps with limit TPs. The exit geometry creates NO edge. - Outcome mix BTC long: stopped 35.1%; rungs hit 0/1/2/3 = 52.9/23.4/10.6/13.1%. - Oracle (always picks the better side) = +53 bps BTC, +68 SOL, +64 DOGE. Anti-oracle mirrors it. - BREAKEVEN DIRECTION ACCURACY: 61.4% BTC / 58.5% SOL / 59.4% DOGE all-taker; 57.7/55.6/56.3% with limit TPs. Repo's best measured OOS direction accuracy is 54.3% (Tier-3 v0). LIQUIDATION CATCH: at a flat 100x with BTCUSDT tier-1 mmr 0.4%, liquidation sits at -0.6% price (-60% of margin), INSIDE the -0.8% (-80%) stop — so at true 100x the stop can never fire on a fresh position; liquidation closes it first. An 80%-of-margin stop only becomes reachable below ~50x, or after a rung is banked and realized profit pushes liquidation out. Any honest test must model liquidation, not just the stop. CAPABILITY GAP: - Backtest: engine/core.py resolves ONE stop + ONE target intrabar and closes the FULL position (_resolve_intrabar). No partial scale-out. Signal.intent already carries add/close_all for LadderEngine (scale-IN); a reduce/rung-aware exit is the missing piece. No leverage/liquidation in the single-instrument engine. - Paper: closer. mean_reversion_adapter already has order_type=limit, market=perp, maker/taker fee split, trail_bps, max_hold_minutes; account 7 proves limit-on-perp works live at 1m cadence. Needs TP-ladder + SL params + margin/liquidation (wallets.py has the primitives). Scripts: /private/tmp/claude-501/-Users-quanvo-Documents-git-repos-personal/2e15dc2b-acb8-4064-aeb2-293ab4d38a87/scratchpad/ladder_null.py and ladder_skill.py. Related cards: #198 (paper engine), #45 (100x leverage research), #73 (minute-level go/no-go).
2 weeks ago
RISE-15560 — AI Autofill from Network Profile: entry point + 3-step modalOrchestrated-feature-dev run for RISE-15560 "Select AI Autofill from Network Profile" (Story, BACKLOG, epic RISE-15559, DF `release_autofill_assessment_AI_phase3`). Repos: rs-backend + rs-frontend. SCOPE: the entry point and the 3-step modal only. AC 1 adds "From network profile" above the 2 existing options on the "Auto-fill assessment" button, shown when the executor org has a Network Profile. AC 2 opens the modal directly when executor org = assessed org; AC 3 interposes a warning dialog when they differ. AC 4 is the modal: step 1 picks data-only (default) vs data-and-documents, step 2 runs the AI analysis and blocks advancing if nothing was found, step 3 applies and reports "X questions answered" per RISE-15252. NOT in scope: RISE-15561 (AI insight badge / section attribution), RISE-15562 (Mixpanel), RISE-15652 (Epic 2 certificate picker). BUILDS ON SPIKE RISE-15634 (card #229, staled). Its decisions S1-S13 are locked inputs, mirrored into tmp/RISE-15560/CONTEXT.md: serialize the NP into 5 section JSON files in GCS and feed them to `rise-document-intake` as documents (no Data Science change needed); read all 5 sections from CDC `cdc_passport_be` and extend the sink, with the live NP API explicitly rejected; copy NP files into the RSC bucket rather than requesting a cross-bucket grant; exclude rt=9 and rt=12 question types. Feasibility is proven — 18/18 answerable questions correct at confidence 1.0 on the real staging service — and the ticket's embedded dev note ("check if we are able to do Step 1") is already answered YES. TWO BLOCKERS on AC 1, both asked by Will in a Jira comment on 2026-09-10 with no reply since: (1) what counts as "has a Network Profile" — a row exists, or a row with data in it; (2) whether the phase-1 and phase-2 dark features are also required to show the option, or whether phase-3 alone is enough. The second appears nowhere in the spike docs. KNOWN INFRA GAP: `ecosystem_organization_document`, the four capacity tables, and the remaining capability child tables are not in the CDC sink. Each addition is Infra-gated and backfill is Infra-only. Certifications, Documents and Capacity sections cannot be fully populated until they land.
2 weeks ago
Infra release 20260922.2 to testnet — add the points-settlement migration and belief-points fee configupredict-infra PR #784 "Upgrading testnet to tag 20260922.2" (branch deploy-20260922.2-to-testnet, base main) currently changes only ci/configs/testnet/pipeline-config.json. Tag 20260922.2 = backend 4c5c1f5b [UBET-4338] (#790), which sits on top of 3fe94b96 [UBET-4337] (#785) — testnet is on 20260922.1 = 5c38cca9, so this release carries BOTH settlement and the surface sweep. Two additions needed: 1. MIGRATION — 20260915000000_add_points_settlement.sql is staged in upredict-backend/migrations/ but absent from Terraform/testnet/postgres/migrations/ (testnet stops at 20260913000000_add_market_bet_user_id_and_points_bets.sql). It drops NOT NULL on market_result.commitment and adds five point_ledger_source_type_enum values (BetPayout, BetRefund, CreatorFee, ReferralFee, TreasuryFee). Must re-run `atlas migrate hash --dir "file://."` after copying. UBET-4338 itself adds NO migration — its only schema change is a view function, which ships with the wholesale sql_schemas deploy. 2. CONFIG — config/testnet/belief_locker/belief_locker-config.yaml already has "beliefPointsConfig": {"enabled": true} on main and on the release branch. Missing are the fee rates. User's call: "same as money market". Money fees are not config at all — they are stamped per-bet by the contract and read off the PlacedBet event, so the rates were read from the testnet DB: the live band since 2025-12-10 is creator_fee_percentage_decimal 10000, operator_fee_percentage_decimal 4000 (earlier bands 1000/4000 and 0/0). Operator fee is the money-side equivalent of the points platformFeePercentageDecimal. treasuryAddress to be omitted: its default is the money chains' shared defaultReferralReceiver, which is 0xbB1a40cE9a9570af2137C7f6489178689A4736c7 on both testnet chains. Known consequence, already accepted by the user (enabled:true was set by them before this session): the testnet frontend has not shipped its belief-points side, so the chain switcher will show a blank-icon entry and an anonymous user selecting it can hit a render crash in the bet-amount input (BigInt(null) on maxBetSizeWei).
2 weeks ago
macOS 27: repeated internet loss + Wi-Fi drops — NordVPN proxy vs iCloud Private RelayDiagnostic investigation on Quan's MacBook Pro (macOS 27.0, build 26A428, arm64). NOT a code change — machine/network troubleshooting, findings recorded for reuse. Session 2026-09-22. REPORTED SYMPTOMS (three, initially believed unrelated): 1. "Bartender 7 kills my internet; force-quitting it restores connectivity." 2. Internet dies completely; only a reboot fixes it (3 reboots on 2026-09-22: 08:23, 10:18, plus earlier). 3. Wi-Fi ("Trung Hung 1") keeps disconnecting; forced onto iPhone hotspot; a second machine on the same AP is fine. ROOT CAUSE (symptoms 1 + 2) — two traffic interceptors fighting: - NordVPN registers a TRANSPARENT PROXY network extension: NESMTransparentProxySession[Primary Tunnel:NordVPN protection], com.nordvpn.macos.Shield (com.apple.networkextension.app-proxy). EVERY socket on the machine carries "flow divert" (15,898 refs in a 14-min window). - iCloud Private Relay is ALSO enabled (com.apple.networkserviceproxy holds a 120KB active NSPConfiguration). It tunnels via mask.icloud.com. - 83% of Private Relay connections (29,088 of 35,019) carry "flow divert" — Apple's tunnel is being intercepted by NordVPN's proxy. - NordVPN's Shield extension constantly fails keychain access: "Failed to talk to secd after 4 attempts" — 3,092 times in 39 min (~20-24/min steady). Independent NordVPN defect. - When NordVPN's side breaks (errno 50 ENETDOWN, "Connection failed to connect 1:50"), Private Relay cannot reach its gateway -> "mask.icloud.com:443 failed resolver" -> infinite retry. - Those retries are issued BY mDNSResponder, so the DNS resolver saturates itself and ALL name resolution dies = "Wi-Fi connected, no internet". Does not self-recover; only a reboot clears it. EVIDENCE (mDNSResponder log lines/min; normal ~1,500): - Before the 10:18 reboot: 10:04=11,425; dead-silent 10:10-10:12 (resolver hung); 10:14=30,961; 10:15=34,892; 10:16=39,015; 10:17=38,893. Reboot 10:18:57. - Before the 08:23 reboot: identical shape — 08:10=26,781, then 10,054/14,180/15,938 up to the restart. - "Failed to talk to secd" spiked in lockstep (73 and 69/min at 08:10-08:11; 137/min at 08:20). - Both shutdowns HUNG and were force-killed by the watchdog (shutdownStall reports), consistent with a wedged network extension. BARTENDER IS NOT THE CAUSE (symptom 1 explained): - /Applications/Bartender 6.app is actually v7.0.4, notarized, signed by Bartender App LLC (24J875RH8J) — genuine, not tampered. - It has 13 entitlements, NONE network-related (no VPN / content-filter / app-proxy / packet-tunnel). It cannot touch traffic. - It IS a trigger: Bartender launched/exited 6x in 10 min with 3 distinct LaunchServices bundle records + a Sparkle self-update. Each re-registration fires nehelper "apps installed" -> nesessionmanager RESTARTS the NordVPN proxy (12 restarts in the window) -> every diverted flow is torn down. Flow teardowns per 10s: 1,062 and 512 at restarts vs 70 baseline. - Duplicate bundles registered: /Applications/Bartender 6.app, stale /Applications/Bartender 7.app, ~/.Trash/Bartender 6.app + 7.app, /Volumes/Bartender 6 + 7 (unmounted DMGs). SEPARATE ISSUE (symptom 3) — Wi-Fi drops have their own causes: - Signal was never weak: RSSI -38 to -46 dBm throughout. - 6 link-downs 15:50-16:38. Two (15:50:47, 15:56:58) coincide exactly with DarkWake from 'Clamshell Sleep'/'Maintenance Sleep' (pmset confirms; 10 Maintenance + 4 Clamshell sleeps that day). Firmware reason: "link down due to beacon loss", state NET_MANAGER_STATE_SLEEP, trigger=dark_wake. - After waking it re-associated to 2.4GHz channel 1 @20MHz instead of 5GHz channel 36 @80MHz. Channel 1 is saturated: crsglitch 2,804,798 (vs 4,401 on ch149); 47 own beacons vs 44 other-BSS. - AWDL (AirDrop/Continuity) held the radio in a permanent real-time schedule: 3,258 "RTG: Active / UserTriggered" events, dwell split ~50/50 between ch1 and ch149. Missed beacons: 18,904 from AWDL, 6,608 home-channel. WHY IT WOULD NOT REJOIN (the decisive one): - 13:36:50 macOS ran CONFIRM BROKEN BACKHAUL PROBE -> TIMED OUT after 3.2s -> concluded the Wi-Fi had no working internet. - 16:41:58 auto-join refused: "Known network profile with recently (<3600s) broken backhaul not allowed when already associated to PH" (PH = Personal Hotspot). macOS deliberately pinned the Mac to the iPhone. - Plus 684 "DeferredTKIP" refusals since 16:00 — the router runs TKIP, which is the "Weak Security" warning, and macOS actively deprioritises TKIP networks. - LIKELY LINK: the 13:36 backhaul probe failed while the NordVPN/Private Relay DNS storm was active. The second machine never got that verdict because it does not run that pair. UNVERIFIED / LIMITS: - SSIDs are redacted in the unified log and /Library/Preferences/com.apple.wifi.known-networks.plist is root-only, so networks were matched by channel+security, not by name. `sudo wdutil info` would confirm. - The total DNS blackout was inferred from teardown/retry spikes, not directly observed (DNS still resolved at 10:56). - Not read: the two shutdownStall reports (user is in _analyticsusers, so they are readable if wanted).
2 weeks ago
Quant: carry's headline result is not reproducible — audit and re-establish it before any buildA 98-agent workflow (survey -> 6 proposal lenses -> 3 adversarial reviewers each -> synthesis) audited the one arm this program still believes in. I then verified the load-bearing claims against the code myself. THE HEADLINE PROBLEM — "always-on carry beats B&H on BTC/ETH/LINK by 1.10/1.30/1.45 Sharpe" is NOT REPRODUCIBLE BY ANY RUNNER IN THE REPO TODAY. - It traces to exactly one place: docs/research/carry-hedge-drag/01-result.md:50, from a single run on 2026-08-05 (baseline commit 33c9409, measurement code 071b6e2). - VERIFIED: carry_verdict.py:272 passes always_on_sharpe into compute_verdict's BENCHMARK slot, and :444 writes bah_sharpe=result.always_on_sharpe. So that runner's question is "does the funding GATE beat always-on carry?" (answer: FAIL, gated 1.198 < always-on 2.314). It structurally cannot grade always-on against buy-and-hold. - VERIFIED: neither src/quant/validation/rehedge_drag.py nor scripts/run_carry_verdict.py contains any buy-and-hold code at all. - NOT a hidden defect: docs/research/carry-hedge-drag/00-pre-registration.md:56 states it outright — "Carry has never been benchmarked against Buy & Hold ... This is deliberate and documented as D17's Step 10/11 encoding, not an oversight." The workflow framed this as a discovery; it is a KNOWN, DOCUMENTED gap. What is genuinely new is that the number published to close that gap cannot be re-derived from current code. MY OWN ERROR: I repeated "carry is the only arm that ever beat buy-and-hold" many times this session as settled fact. It rests on one non-reproducible document. Corrected to the operator. THE NUMBER'S ERROR BAR IS MOSTLY BOOKKEEPING (per the workflow): the identical always-on arm reads -0.249 (bare test slice via _mean_baseline_sharpe), +1.20 (per-symbol non-overlapping windows), +2.31 (warm-sliced), +5.50 (zero cost). A ~6-Sharpe spread that is 100% convention and 0% market. TWO DEFECTS I VERIFIED STRUCTURALLY IN THE CODE: 1. GROSS RAMP — engine/carry.py:104-111 sizes BOTH legs once at entry (notional = gross_per_leg * equity_open) and never resizes; the comment at :99 says matched units deliberately avoid re-trading every bar. So notional/equity drifts freely with price and equity. Agents measured BTC 0.503 -> median 1.351 -> max 2.731 (total gross 5.46x against a declared 1.00x cap; ETH 10.23x; SOL 31.16x), and de-levering flips SOL 1.348 -> -0.168 and moves BTC 4.232 -> 3.790. MAGNITUDES NOT INDEPENDENTLY VERIFIED BY ME — the structural cause is confirmed, the numbers are agent-measured via scratchpad scripts (measure_gross_ramp.py, measure_collateral.py). 2. F5 ENTRY FILL — always-on sets pending_on at bar 0 and fills at bar 1 of the WIDENED panel, inside the warm prefix carry_verdict then slices off, so its entry cost never lands in reported returns. Repo's own estimate: +0.4 to +1.0 Sharpe per fold. Documented at docs/research/structural-edges-audit/flaws-to-fix.md:76-82. NOTE: that file lists F5 among items "fixed via meaningful-red TDD", but the fix was to RELABEL 2.314 as "frictionless", not to charge the fill — the behaviour still stands. 3. COST NEVER PRICED FOR CARRY — run_carry_verdict.py hardcodes CostModel() with no fee flags, CarryEngine charges ONE CostModel to BOTH legs, run_multi_carry_verdict.py exposes no cost knobs. No carry verdict has ever been priced at the verified futures-taker schedule (spot 10 / futures 5 bps). Agent reconstructions (UNVERIFIED, hypothesis only) put the 9-leg basket at -0.575 -> +0.93..+1.12 and the mean-of-9 baseline at -0.249 -> +0.91 at 14 bps, with the diversification lift turning positive — which would reopen the basket refutation on cost grounds. DSR IS THE WRONG INSTRUMENT HERE: n_trials is hardcoded to 1, deflation_applied needs >=5, and B&H itself scores DSR 0.0000 at n_trials>=10 at 1h. Use a paired block bootstrap of the DIFFERENCE (carry minus B&H on the common index), block length swept at funding-regime scale (336/720/1440 bars, not the ML default 24), and declare DSR inapplicable with the dated_carry_verdict precedent. Publish block-length sensitivity, not one number. HONESTY NOTE worth carrying into any write-up: in a coin-matched delta-neutral book the spot gain IS the perp mark loss, so a 1x liquidation forfeits ~25 bps of equity and leaves a naked spot leg, not a wipeout. 1x breach count on BTC/ETH/LINK is 0/1/1 folds out of 109 — a caveat, not a retraction. At 3x it is much worse (see card #198's liquidation table). Full synthesis: /private/tmp/claude-501/-Users-quanvo-Documents-git-repos-personal/2e15dc2b-acb8-4064-aeb2-293ab4d38a87/tasks/wj39dd7cd.output (98/98 agents, 8.7M subagent tokens). Related: #258 (ladder/minute-frame refutation), #198 (paper-trading ops audit).
2 weeks ago
UBET-4345 AI comment moderator — quality gate for points + gibberish filterJira UBET-4345 "AI comment moderator" (Task, Todo, Sprint?, reporter Daniel Jiwoong Im, assignee Quan Vo, parent epic UBET-3691 "BM - Our Own Comments"). Ticket body is 3 lines: (1) decide whether to give points or not, (2) if it is not a word but random characters don't allow to display, (3) judgment criteria — meaningful? / gets conversation going? / super emotionally stimulating? CURRENT STATE (BE, read pass done): OpenAI is the only LLM vendor in belief_locker. Comment path = synchronous OpenAI Moderation API (omni-moderation-latest) in commentLlmService.ts, called from routes/commentOnMarket.ts before any write; flagged categories (default hate, self-harm) reject with MARKET_TEXT_MODERATED 400. Fail-open on any API error. Market path = marketLlmService.ts does moderation + gpt-4.1-mini translation (JSON schema) + text-embedding-3-small embedding, sync at create and via FetchTranslationsTask. Images = AWS Rekognition DetectModerationLabels. Belief points for comments are decided by pure SQL (appSql/get_comment_ledger_decision.sql) on structure only (first reply per author/market, first credit-worthy interaction per market) — no content signal. market_comment has NO hidden/moderated/score column, so there is no way today to store a comment but suppress its display. Running the orchestrated-feature-dev pipeline; workspace upredict-backend/tmp/UBET-4345/. Sibling card #194 (UBET-4241 comment sorting) is the same epic but a different ticket.
2 weeks ago
Quant: measurement-stack defects — DSR var_trials input, carry fold geometry, corrupt LINK barFrom a 74-agent workflow on legitimate ML/TA uses. The headline was not an ML finding — it was that the measurement stack is mis-reporting things. I VERIFIED the two most consequential claims against the code and data myself; the rest are agent-measured and flagged as such. 1) DSR var_trials IS FED THE WRONG VARIANCE (verified by me, structural): - carry_verdict.py:263, multi_carry_verdict.py:374, ladder_verdict.py:411, ml/prediction/walk_forward_ml.py:341 all pass statistics.pvariance(fold_sharpes)/bpy — the fold-to-fold variance of ONE strategy. - sweep.py:116 passes statistics.pvariance(result.all_oos_sharpes)/bpy — variance ACROSS TRIALS, which is what Bailey/Lopez de Prado's expected-maximum term wants. - walk_forward_ml.py:339 has a comment claiming it mirrors sweep.py. It does NOT — fold_sharpes (one config, many folds) is not all_oos_sharpes (many configs). That false equivalence is probably how this survived review. - WHY IT MATTERS: fold-to-fold Sharpe variance for a single strategy is large, so feeding it as across-trial variance inflates the expected-max hurdle and crushes DSR toward zero. This is consistent with the recorded "B&H itself scores DSR 0.0000 at n_trials>=10 at 1h" symptom — i.e. the DSR-unfireable finding may be an ARTIFACT OF THIS INPUT, not a property of the 1h interval. - Agent re-run (NOT verified by me): with across-trial variance the gate grades normally at 1h — DSR 0.914 / 0.727 / 0.096 at n_trials=32 for across-trial sd 0.3 / 0.5 / 1.0. - BLAST RADIUS: every DSR-gated verdict in this program was read through this input. The z-score ladder arm was failed on DSR 0.930 <= 0.95. Re-read all of them after the fix. - Memory quant-trading-dsr-gate-unfireable.md has been annotated so a future session does not treat it as settled. 2) CORRUPT BAR IN THE STORE (verified by me, exact): - linkusdt 2020-03-12 10:48:00 UTC has low = $0.0001 against open 2.1372 / close 2.20. Exactly one such bar across LINK; SOL has none by the same screen (low<50% or high>200% of close). - engine/core.py _intrabar_hits tests `bar.low <= stop` for a long, so EVERY LINK backtest carrying a long stop has been spuriously stopped out on that bar, and the triple-barrier labeller touches the same low. - Agents also claim SOL's carry leg shows 67 bars above 100 bps and a max single-bar move of 1,459.9 bps — that is a BASIS/desync claim, not a raw OHLC anomaly, and my screen would not catch it. UNVERIFIED. - Fix: an outlier guard in data/store.py rejecting a bar whose log-range exceeds a large multiple of its trailing median. 3) CARRY FOLD GEOMETRY CHARGES A ROUND TRIP A LIVE HOLDER NEVER PAYS (agent-measured, NOT verified by me): - multi_carry_verdict.py:191 runs each fold on the bare test_idx with fresh cash and liquidate_at_end=True, so all 18 legs open and close every 500 bars, 95 times. - Claimed size: ~24 bps/fold of forced churn against ~26.9 bps/fold of carry income — the harness cost is ~100% of the strategy's revenue. Same BTC arm reads +2.31 at test=500 and ~+5.10 at test=8000 with nothing else changed. - Combined with F5 (always-on's entry fill lands in the discarded warm prefix and is free, while the gated arm pays every toggle inside the scored window), the two carry arms are scored under DIFFERENT cost regimes. See card #264. 4) CARRY HAS NO MARGIN MODEL (agent-measured, NOT verified by me): - CarryEngine's only guard is equity <= 0, which a delta-neutral book can never trip. paper/wallets.py:50 already has liquidation_price = entry * (1 - mmr + 1/leverage) and dated_adapter.py:263 already enforces it. - Claimed: 12/97 SOL folds, 8/109 LINK, 3/109 ETH, 1/108 BTC cross the liquidation price today and still report a finite Sharpe. At a fixed 1.0x gross the book IS fundable with zero in-sample liquidations (max perp-leg leverage BTC 2.61x, SOL 1.23x, AVAX 1.11x). Repo DEFAULT_LEVERAGE = 3.0 liquidates on a +32.83% rally, which BTC/SOL/AVAX all exceeded — consistent with my own independent measurement on card #198. - Also: CarryConfig.__post_init__'s 2*gross_per_leg <= 1.0 check is a construction-time assert on a config field; the live book allegedly breaches it by up to 27x on SOL. Should be a runtime invariant in the engine loop. 5) THE ONE REAL ML/TA WIN (agent-measured, NOT verified by me): Parkinson range volatility sqrt(mean(ln(high/low)^2)/(4 ln 2)) beats the incumbent close-to-close pstdev on Spearman 9/9 AND QLIKE 9/9 across the 9 coins at 1h (BTC rho 0.6068 vs 0.5816; QLIKE 19% lower). An EWMA(0.94) control wins Spearman 9/9 but LOSES QLIKE 8/9 — so the gain is the RANGE, not the weighting. Information argument (high/low are order statistics the close discards; Parkinson 1980 ~5x efficiency), not a pattern claim. Path to P&L is indirect: it sharpens every sigma-consuming component. ORGANISING RULE the workflow extracted, worth keeping: every construct that survived on this data predicts a SECOND MOMENT or a mechanical identity; every construct that died predicts a SIGN. Second moments are conserved and observable; signs are competed away. Full output: /private/tmp/claude-501/-Users-quanvo-Documents-git-repos-personal/2e15dc2b-acb8-4064-aeb2-293ab4d38a87/tasks/wemf5ao6x.output (74/74 agents, 7.5M subagent tokens). Related: #264 (carry credibility), #258 (minute-frame refutation), #198 (ops).
2 weeks ago
Quant: order-book depth + cascade absorption — free data exists, but the feed lies during cascadesOperator proposed using bid volume as a signal, tied to cascade absorption. 79-agent workflow (6 survey angles -> 3 adversaries per candidate -> plan). 78/79 agents, one TLS drop on a candidate that still got 2 of 3 votes. I verified the data-availability claims myself. HIS IDEA IS GENUINELY NEW INFORMATION, NOT A REPEAT. Spearman(depth imbalance, taker_buy ratio) = -0.025 on 519,433 joined BTC minutes. taker_buy_base_volume (aggressor flow) is already a feature in every failed direction run (config.py:93/98/132-134); resting depth is a different quantity. DATA AVAILABILITY — I VERIFIED THESE AGAINST THE LIVE HOST: - futures/um/daily/bookDepth: 2023-01-01 -> 2026-09-21, LIVE. ~470KB/day, 0.59GB BTC full history, ~5.6GB for 9 coins. (My first listing hit the 2000-key cap and looked like it ended 2024-05-17; paginating properly shows it is current. Watch that trap.) - futures/um/daily/bookTicker: DEAD ARCHIVE — exactly 320 files, 2023-05-16 to 2024-03-30, then Binance stopped. 53GB for BTC alone. Do not build on it. - futures/um/daily/liquidationSnapshot: DOES NOT EXIST for USDⓈ-M (zero keys). Exists only for COIN-M and is discontinued. There is NO forced-liquidation ground truth for the market he trades. - futures/um/daily/metrics: OI + 3 long/short ratios at 5-min cadence, BTC from 2020-09-01 (others 2021-12), 0.19GB for all nine. Cheapest dataset, and the one that bears on carry timing. - futures/um/daily/aggTrades: 2019-12-31 -> now, ~18MB/day, 44.7GB/symbol full, ~3.6GB if only cascade days. - Ingestion measured on this machine: 120 bookDepth files in 12.3s at 10-way parallel; 2,736 metrics files in 30.5s. Bandwidth is not the constraint. binance_dump.py needs a _days() generator + daily loaders + a long->wide pivot; store.py read_store hardcodes OHLCV + RAW_EXTRA. WHY DEPTH DOES NOT DELIVER (agent-measured, not verified by me): - It is NOT an order book: 10 rows per snapshot, every ~30s, CUMULATIVE notional within +/-1/2/3/4/5% of mid. No price levels, no queue, no top-of-book. At BTC $110k the tightest band is $1,100 wide. The +/-0.2% near-touch band only appears ~2026-01-15, so near-touch designs have 5.5 months of single-regime data. - GEOMETRIC TRAP: the band is measured off the CURRENT mid, so a price move slides the window along a static book. Spearman(bid-depth change, 5m return) = -0.53, ask = +0.51 — near-perfect mirrors. Bid depth RISES as price falls. A naive study measures a coordinate system, not demand. - Economically hopeless at 30s-60m: raw depth-imbalance IC +0.012/+0.027/+0.042/+0.049 at 1/5/15/60m, collapsing to +0.008-0.013 after controlling for past returns — ~80% is restated short-horizon mean reversion, already killed by the 36-config TA sweep. Decile spread of forward 60m return +2.34 bps, non-monotone, vs a 10-12 bps round trip. Causal form (rolling z, |z|>1, h=60m, n=183,087): +0.57 bps gross. - Adds nothing to vol forecasting: past-60m RV predicts forward-60m RV at +0.79; depth's incremental contribution after residualizing is -0.012. - Cost model correction runs AGAINST us: live walk-the-book at $50k gives BTC 0.006 bps, ETH 0.018, SOL 0.43, DOGE 1.3-2.1, LINK 3.2-4.9, AVAX 3.5-5.1 vs the 2.0 bps/leg assumed in costs.py. At $200k LINK/AVAX are 9-12 bps/leg. Sizing constraint on alt baskets. - CAPACITY IS A NON-ISSUE: the +/-1% bid band holds $17.9M at its 1st-percentile thinnest and ~$120-190M typically. A $200k clip is 0.1-1% of the thinnest book. THE FEED LIES EXACTLY WHEN IT MATTERS — the single most important finding here: - On 2025-10-10 (largest liquidation cascade on record) the bid series PINNED at exactly $2,142,547.09 / 17.652 BTC for 379 consecutive snapshots = 189 minutes, straight through the 21:19 crash, while the ask side updated normally. - A frozen LOW value reads as a maximal liquidity vacuum — THE ARTIFACT CONFIRMS THE HYPOTHESIS. Multiple independent passes found "77x liquidity withdrawal" and it was a stalled feed. - Not a one-day problem: 10.67% of 2025 BTC snapshots repeat the previous value to the cent; 2025-04-16 -> 2025-05-18 (34 days) has exactly one distinct value per band per day. - MANDATORY LOADER RULES: drop runs of >=3 identical values before anything else; parse `percentage` as Float64 (written "-5" before ~2026-01 and "-5.00" after, which raises a polars ComputeError on Int8). CASCADE ABSORPTION — latency is genuinely fine, adverse selection is brutal: - Latency NOT the problem (measured on real aggTrades 2025-10-10 and 2024-04-13): price dwelt below a -5 bps limit for 7 to 300 SECONDS with $0.33M-$1.8B of resting-bid-hit notional below it. 60-90ms is 2-4 orders of magnitude inside that. - Aggressor imbalance does NOT detect a cascade: through the 2025-10-10 peak aggressive-sell share ran 50-60% vs a 0.501 day mean — price fell 5% in a minute on roughly BALANCED flow. What spikes is INTENSITY: prints/sec 6 -> 1,000-3,000, makers consumed per sweep 8 -> 127. Cascades are violence, not one-sidedness. - SITTING ON THE BID IS WORSE THAN TAKING: limit 0.5% under the signal close returns +58.7 bps at H60; plain market buy +62.3; events where the limit NEVER FILLED returned +137.4. You get filled precisely when the flush continues. Adverse selection, quantified. - WINDOW OVERLAP is the largest error in every cascade number so far: at a 0.5% wick BTC goes from +52 bps (t=6.4) to +7 bps (t=1.0) once de-overlapped to one position per 60 minutes. - REPORTING MEDIANS is the second largest: at a 10 bps offset over 2,241 fills the median is +5.83 bps and the MEAN is -1.36; removing the 5 worst days flips it to +0.98 — and those 5 days (2021-05-19, 2020-08-02, 2020-05-10, 2021-04-18, 2021-09-07) ARE the actual cascades. You buy ordinary dips and get run over by the real ones. - A STOP-LOSS DESTROYS IT: -5%/H60 goes from +121.3 to +42.2 bps gross and 66% -> 44% win rate with a 3% stop. The edge requires holding through adverse excursion => no meaningful leverage. - SURVIVORSHIP IS HORIZON-DEPENDENT: alive-9 vs dead-4 mean forward return H15 +27.8/+19.6, H60 +46.9/+36.2, H240 +78.3/-16.2, H480 +99.7/-30.4. The short-horizon effect survives; the LONG-horizon effect is ENTIRELY survivorship. - LUNA disproves the "liquidity filter saves you" defence: on 2022-05-09, the day before the death spiral, LUNAUSDT was the 3rd most liquid USDT perp by trailing-30d dollar volume, ahead of SOL/XRP/DOGE/BNB. Any live filter includes it. Worst name in the study at -242 bps mean, -71% worst event. - LUNA and FTT PREDATE bookDepth (May/Nov 2022 vs a 2023-01-01 start), so depth data structurally CANNOT see the two observations that matter most. Run the survivorship bound on klines. - The true delisting cohort is 31 USDT perps (most of the apparent 134 are BUSD retirements or renames like MATIC->POL); 26 are already downloadable. data.binance.vision does not delete delisted symbols. BUILD ORDER (inverted on purpose — do not download first): 1. ZERO-DOWNLOAD precondition test, an afternoon, no bytes: ungated cascade fade (-3%/-5% 15-min drop) on the perp klines already on disk plus the 26 delisted symbols, de-overlapped to one position per horizon, entry at next bar open, H in {15,60}, 2023-01-01 onward. Report the MEAN (never the median), the worst single event, and a day-block bootstrap over distinct calendar days (8,294 trades sit on only 995 days; 25% fire in hours where >=7 of 9 coins move together). KILL: if the 95% day-block CI net of 10 bps does not exclude zero, the family is closed and no order-book data can rescue it. EXPECT IT TO FAIL — the best existing measurement is gross +19.26 bps, day-block CI [+5.5,+33.3], net at 10 bps [-4.5,+23.3], already a failure. 2. Only if (1) passes: aggTrades on ~300 cascade days (~3.6GB, ~12 min) to price the FILL. The entry minute's median high-low range is 166 bps and the whole claimed edge is 14-50 bps, so where inside that minute the order lands decides everything. A mild adverse-fill assumption already turns the ungated arm from +14.2 to -17.9 bps. KILL: realized one-way cost on DOGE/LINK above ~6 bps and it is dead at any gate strength. 3. Only then bookDepth, and only as a SELECTOR conditional on (1) and (2) being net-positive — never as a standalone predictor. Mandatory: drop >=3 identical-value runs, Float64 percentage, and residualize every depth change against the contemporaneous return before use. Full output: /private/tmp/claude-501/-Users-quanvo-Documents-git-repos-personal/2e15dc2b-acb8-4064-aeb2-293ab4d38a87/tasks/wg73vqmoc.output (9.1M subagent tokens, 2,233 tool calls, ~15GB downloaded to scratchpad). Related: #266 (measurement-stack defects), #264 (carry credibility), #258, #198.
2 weeks ago

Done

137
#21
P0
Integrate AI-Kanban with OpenClaw (feature idea + OpenClaw setup reference)FEATURE IDEA Integrate AI-Kanban with OpenClaw so the board can be driven from a chat channel (e.g. Telegram): create/query/move cards, get progress pings, and potentially have OpenClaw agents pick up and work cards. OpenClaw already exposes MCP tools + a paired Telegram control channel, so a thin bridge to the ai-kanban-dispatch MCP tools is the likely integration surface. --- OpenClaw local setup (reference, as of 2026-07-06) --- Host: Quan's MacBook Pro (macOS 26.5.2 arm64). OpenClaw 2026.6.11 (brew: /opt/homebrew/bin/openclaw). Config: ~/.openclaw/openclaw.json. Gateway: LaunchAgent, ws://127.0.0.1:18789. GATEWAY: was crash-looping every ~10s ("Gateway start blocked: existing config is missing gateway.mode"). Fixed with `openclaw config set gateway.mode local` + relaunch. Now listening/reachable. doctor --lint: 0 errors (only warning: gateway.auth.token stored plaintext — cosmetic). MODEL / AUTH: `openclaw configure --section model` → provider Anthropic via claude-cli OAuth (mode=oauth, profile anthropic:claude-cli) = REUSES the Claude Code subscription (no metered API key). Default model = anthropic/claude-opus-4-8 (fallbacks 4-7 / sonnet-4-6 / 4-6). Verified with a live agent turn. NOTE: subscription OAuth in a 3rd-party tool is a ToS gray area; fine for light manual use, risky for unattended. Cost on this route = Max-plan quota, not $. Switch models per-conversation from Telegram with /model <id>; /model default resets to Haiku/Opus default. TELEGRAM CHANNEL: bot @quan_vo_open_claw_bot, dmPolicy=pairing. Paired + set as command owner (telegram:1767875031, "Quan Vo") via `openclaw pairing approve telegram <code>` — only this account can command it or switch models. Inbound + outbound + agent auto-reply all confirmed working (first msg was silent due to a one-time auth-profile session reset; resend worked). BROWSER: plugin enabled. Isolated `openclaw` profile (CDP, port 18800) running. `user` profile = existing-session / chrome-mcp attach to real Chrome (where Bitwarden lives) — NOT yet wired (real Chrome not exposing a debug port; chrome-mcp bridge needed). Existing-session profiles also have some unsupported browser actions (ACT_EXISTING_SESSION_UNSUPPORTED). PRIMARY USE CASE BEING BUILT (separate from the ai-kanban integration): on-demand AWS SSO login for Claude Code's aws CLI. Flow: OpenClaw runs `aws sso login --sso-session upredict --no-browser` → navigates the printed verification URL in the real Chrome → Bitwarden autofill login (company SSO portal TTL ~1h, NO MFA) → click Confirm → Allow → token cached → aws works. Trigger = MANUAL, on demand from Telegram ("refresh my AWS login"), not a cron. Model Haiku (cheap). Hard rule: stop and ping on any unexpected page. (upredict SSO covers prod=eu-south-1 + testnet=us-east-2 DevOps — token grants both.) OPEN ITEMS 1. Wire `user` profile bridge so OpenClaw can drive the real Chrome (Bitwarden). 2. Write the "refresh aws" OpenClaw skill (pings back over Telegram). 3. Test AWS flow end-to-end. 4. (This card) design + build the AI-Kanban <-> OpenClaw bridge.
3 months ago
Ladder verdict soundness: stop_z ceiling guard, baseline sizing, true B&H arm + re-run bake-offOpus audit of run #31 found the run did not test its configured strategy. FLAWS: (1) stop_z=5.0 is mathematically unreachable — RollingZScore uses statistics.pstdev over the full lookback window, so max |z| = (n-1)/sqrt(n) = 4.359 at lookback=20; 0 of 18415 trades exited on 'stop', max entry_z observed 4.3537. Combined with max_hold_bars=0 (disabled) the run had NO loss exit and NO time exit. (2) rung 2 (|z|>=4.0) fired 8 times in 15550 baskets — 84% single-entry, so the ladder thesis was barely exercised. (3) ladder_verdict.py:313 calls run_backtest with no size kwarg so the baseline arm runs at 100% of cash vs the ladder's ~33% — arms compared at ~3x different notional. (4) bah_sharpe stores the single-entry MR arm, not real buy-and-hold (run #30 stores true B&H +1.0857 in the same column) so #31 cannot answer 'should I just have held?'. (5) The gate's real comparator max(baseline,0.0) is not persisted, so the record reads as 'beat the baseline but failed'. VERDICT ON #31: the FAIL is trustworthy in direction — frictionless per-trade edge is -15.8bps over 18415 trades and 47000 OOS bars, i.e. it loses at ZERO cost; fees $120,165 vs gross-of-fee PnL -$118,566. But the ladder thesis is UNTESTED, not refuted. REFUTED (my own suspicions, checked and wrong): purge=0 is not leakage (neither arm trains a model; each fold builds a fresh engine over test_idx only). train=2000/test=1000 is BETTER than the old test=120 (47 vs 398 folds, 47000 vs 47760 OOS bars, 940 vs 7960 warm-up bars wasted). The negative-baseline gate is already correctly floored at max(baseline,0.0). No repeat of the run-#28 blowup artifact (equity never exceeds its 10000 start; largest step +3.03%). SCOPE (user-approved): guards + code fixes + re-run as a 2-config bake-off (rescaled rungs at lookback=20 vs original [2,3,4] at lookback=50), realistic costs only — no zero-cost diagnostic.
2 months ago
claude-usage: migrate node:test -> vitest, and add Playwright happy-path e2eFollow-on to card #132. Two tracks, ordered: (A) unify the test runner, then (B) add the browser layer that does not exist at all today. # Track A — migrate the 18 node:test suites to vitest Today the repo runs TWO runners: `node --test tests/*.test.mjs && vitest run`. The `&&` means a CLI failure HIDES the server results entirely. The 18 tests/*.test.mjs files predate card #132 (added in 40b1138 parser-core, 5145067 sync). They cover plain .mjs — src/parser/, bin/sync.mjs, hooks/ — with no TypeScript, no JSX, no @/ aliases, which is why the bare node runner was sufficient. vitest.config.ts currently includes only src/**/*.test.ts. ## Why migrating is safe (the fidelity objection does not apply) vitest transforms modules, and these tests exist partly to verify what real `node` does with a real .mjs file. But the hook/CLI tests `spawn(process.execPath, [hookPath])` — a REAL child process at a REAL file path. The code under test is never transformed by either runner; the runner only affects the harness. Migration costs nothing in fidelity. ## Shape Use a vitest WORKSPACE with two projects, not one merged config — the CLI tests must not inherit mongodb-memory-server's 60s timeouts or MONGOMS_SYSTEM_BINARY. ## Wins - One command, one watch mode, one coverage report; no `&&` masking failures. - Readable diffs on failure (node:assert/strict prints a poor one). - `retry` becomes available for genuinely timing-dependent tests. ## Gotchas to check before committing - Keep node:assert or convert to expect? RECOMMENDATION: keep node:assert in the migration commit so it is a provable pure move; convert opportunistically after. Mixing a runner swap with an assertion rewrite makes any failure ambiguous. - vitest runs files in parallel workers by default. These use mkdtempSync and port 0 so they look safe, but the ones mutating process.env need a real look. - node:test subtests `t.test()` map to `describe`, not 1:1. - Keep `npm run test:cli` / `test:server` working, or update every reference to them. # Track B — Playwright happy-path e2e Everything under src/app/*.ui.tsx has ZERO automated coverage — type-checked by `npx tsc --noEmit` and nothing more. Scope here is the HAPPY PATH, not exhaustive. WHY NOT jsdom: deliberately rejected during card #132. jsdom returns zeros from getBoundingClientRect, so a chart axis tick "survives" in the test while a real browser deletes it — a false green on exactly the claims that need checking. Recorded in docs/features/claude-usage-dashboard-insights/spec.md under "Things a test cannot tell you here". ## Happy-path candidates - Dashboard loads authenticated and renders all six cards without an error boundary. - Range picker: choose a preset, the URL updates, the cards re-read it. - Cost split: switch tab (machine/project/model/repo), chart + totals table both change. - Session list: re-sort, page forward, click into /session/[id]. - The URL is the single source of view state: preset/tab/from/to/sort/page survive a reload and a shared link. ## Three open rendering questions (from the #132 spec) — verify, do not necessarily fix here 1. Today's bar is not marked partial on the cost-per-day chart (only the KPI tile says so). 2. At 90-day and all-time windows the X axis is crowded; no bucketing or tick strategy was built. 3. A fully-unpriced day and a gap-filled empty day both draw a zero-height bar. Distinguishable in the data (eventCount) and the range-level statement fires, but not by looking at one bar. ## Setup constraints - Auth: every dashboard route 307s to /login (proxy.ts). storageState vs a test-only bypass — FIRST DECISION, has security implications (an env-gated auth bypass shipped to prod is a real risk). - Dev-server isolation: an e2e-owned `next dev` shares .next with a running dev server and hangs. Use NEXT_DIST_DIR + tree-kill teardown. - Data: the local mongodb://localhost:27017/claude-usage is READ-ONLY — only copy of the backfill, must never be written to. Either read it read-only or seed a separate fixture DB. - Repo rule: never run `npm run build` / `npm run dev` directly.
2 months ago
claude-usage: session titles, one-command machine setup, chart bucketing, folder-grouped repo viewFollow-on to #142. Four independent features, ordered by what unblocks the operator soonest. # F1 — One-command setup on a second machine (do first: it is blocking real use) Today, adding a machine means: clone, `node scripts/install.mjs`, then HAND-WRITE `.env` with values fetched out of Pulumi. The install script literally ends by telling you to go do that yourself. A client machine needs only TWO values — `CLAUDE_USAGE_API_URL` and `CLAUDE_USAGE_SECRET`. No MongoDB: only the deployed server talks to Atlas. The repo is PUBLIC (github.com/votrungquan1999/claude-usage) so cloning needs no auth. AC: one command installs and configures a new machine end to end, with the secret supplied once (flag or prompt). `.env` is written by the tool, not by hand. Re-running is safe. Note: `resolveDotEnvPath` looks in `~/.claude/claude-usage/.env` first, then the repo root — so writing to the installed path is enough. # F2 — Session titles in the session list The list is currently NOT distinguishable: many rows share the same project and machine, and the session id appears only in the link target. Claude Code already records a name. Transcripts carry `{"type":"ai-title","aiTitle":"Build Claude session usage dashboard","sessionId":"..."}` as its OWN record type — not attached to assistant turns — and it can be rewritten during a session (3 such records in one observed transcript), so LAST ONE WINS. DECIDED, WITH EYES OPEN: this widens the aggregates-only rule. `aiTitle` is model-generated FROM the conversation, so it is content, not an aggregate. The operator explicitly accepted this after being shown the tradeoff. Update the sync spec's allowlist to say so — the guarantee is now "aggregates plus the session title", and the event-mapper allowlist test must be updated deliberately, never silently. DO NOT ever upload `lastPrompt`, which sits in the same transcripts and is raw prompt text. AC: the session list shows a name column; sessions without a title degrade gracefully; existing history needs a backfill re-run to populate. # F3 — Chart bucketing: at most 20 bars At 90-day and all-time windows the X axis is unreadable (see #132 spec, now marked "being fixed"). DECIDED: cap at 20 bars on EVERY window over 20 days — one rule, not two. This visibly changes the 30-day DEFAULT view to ~15 two-day buckets; the operator chose that over leaving 30d untouched. Applies to BOTH per-day charts: cost-per-day AND model-mix, so the two stay readable together and their X axes line up. SUBTLETY: model mix is share-of-spend. Buckets must sum cost and RECOMPUTE the share. Averaging percentages is wrong and will look plausible. Suggested home: `dashboard-format.ts`, alongside the existing `rollUpDailySavings` / `rollUpEfficiencyByDay` helpers. Note it compiles into the CLIENT bundle, so no pricing arithmetic there. # F4 — Repo view groups repo-less spend by FOLDER REVERSES a #132 decision. Today every repo-less row collapses into one `(unattributed)` bucket, on the reasoning that "no repository" is the answer itself and spreading it across project names would hide that it is 54.8% of spend. The operator has now seen that data and wants the opposite: group by folder instead. Real numbers behind the current bucket: - git-repos/personal — 27,177 events, $3,225.04 - git-repos/concrete_engine — 2,319 events, $457.39 - .openclaw/workspace — 649 events, $195.70 Context worth keeping: these are unattributed CORRECTLY. `~/Documents/git-repos/personal` is not a repo — it CONTAINS 40 of them — so a session launched there has no single repository. Attribution is by the session's cwd, not by which repos got edited. AC: the Repo tab shows repo-less spend split by folder rather than as one bucket. Mark the superseded #132 decision outdated and update that spec section — it currently states the opposite invariant as deliberate.
2 months ago
claude-usage dashboard: drill down from a project/repo row to its sessionsOperator ask (2026-08-05, with screenshot of the Cost-per-day card on the Project tab): clicking a project or repo — in the totals table below the chart — should navigate to a detail view showing the SESSIONS for that project/repo. Today nothing on that card is clickable; the only way into a session is the flat "Sessions in this range" list or pasting an id into the lookup form. ## Constraints already established (do not re-derive) 1. **A Repo row spans several projectSlugs.** `mergeRepoRows` groups by `repoKey` (SHA-256 of the normalised git remote) and labels the group with the SHORTEST projectSlug in it. `mergeProjectRowsByRepoKey` does the same on the Project tab, minus the repo-less collapse. So "sessions for this row" is a group of slugs, not one slug — the four known worktree pairs (AI-rules-repo, AI-Kanban, bas-attendance, personal-infra) all merge this way. 2. **`repoKey` must never reach the browser** (#132 D31 — it is dictionary-confirmable, an identifier rather than an opaque token). The link therefore carries the visible LABEL; the server resolves label -> repoKey -> the slug set. Not a design choice. 3. **`(unattributed)` is a real, clickable answer** on the Repo tab — now a genuine ~7% residue of work outside any repository, post per-turn attribution (#143 F5). Drilling into it means "sessions with no repoKey", which is a different query shape from a named repo. 4. Session ownership is per EVENT, not per session — per-turn attribution (#143 F5) means one session's events can carry different projectSlug/repoKey values. So "sessions for a repo" must mean "sessions with at least one event attributed to it", and any cost shown must be the cost attributed to THAT repo, not the session's whole cost. This is the trap in the whole feature. ## Out of scope unless asked Making the chart bars or legend clickable (recharts internals); the Machine/Model tabs, pending the operator's answer.
2 months ago
review-changes: label each finding's origin (introduced vs pre-existing)Add a finding-level `Origin` field to the review-changes skill so a reviewer can tell whether the change CAUSED the problem or it was already there. The agent surfaces origin; it must not judge/discount on it (severity and confidence stay independent). Audit of the current skill (all 3 variants): no provenance field exists, and pre-existing issues are hard-dropped at 4 chokepoints — lens-common.md:7 (Scope) and :15 (What NOT to flag), node-verify.md:26+33 (pre-existing → REFUTE), node-merge.md:29 (confidence 0–25) + :57 (drop list). Also node-lens-correctness.md:19 and node-lens-performance.md:14. Three cases the current binary rule gets wrong: 1. Touched-but-not-caused — diff moved/reformatted a line that was already wrong; surfaces unlabeled and mis-blames the author. 2. Pre-existing-but-newly-reachable — faulty code outside the diff, but the change is what now calls/exposes it. Currently REFUTED as out of scope = real false negative. 3. Pre-existing-and-worsened — existing N+1 the diff now runs per-request. Same deletion. USER DECISIONS (2026-08-06): - Scope = "Label + unblock newly-reached": three Origin values (introduced / pre-existing — touched / pre-existing — newly reached), AND verify stops auto-refuting the newly-reached+worsened case. Unrelated pre-existing issues stay OUT of the report. - Effort = "Free signal, escalate when unsure": infer origin from the patch (+ line vs context); when ambiguous (moved code, rename, refactor) the lens marks it unconfirmed via Needs verification and the verify phase resolves it with git blame/log. No blanket git blame per finding. Blast radius: skills/claude-code/review-changes (SKILL + 10 nodes), skills/cursor/review-changes (SKILL + 9 nodes, merge inline in SKILL.md), skills/antigravity/review-changes (single 213-line SKILL.md). Do NOT hand-edit .claude/skills/ dogfood copies — CLI-regenerated.
2 months ago
claude-usage: sync staleness signal, per-session carry timeline, dashboard heading structureThree independent features, ordered by value. Consolidates cards #158/#159/#160 (archived in favour of this one). Chosen by the operator from a wider list on 2026-08-08. # F1 — Surface when a machine last synced (do first) Every failure mode in `bin/sync.mjs` degrades to a silent no-op **by design** — no network, non-2xx, missing `.env`, unreadable transcript all return `{sent: 0}` and throw nothing. There is no local watermark, no `lastSynced` field, and no dashboard surface showing when a machine last succeeded. A machine can stop uploading entirely and nothing anywhere says so. Not hypothetical: 2026-08-04 to 2026-08-06 this machine sent nothing — first `CLAUDE_USAGE_API_URL` pointing at a dead `localhost:3001`, then a stale `CLAUDE_USAGE_SECRET` returning 401 underneath it. Both present identically to "nothing to send". Caught only because the operator looked at the chart and noticed the bars had stopped advancing. **AC:** a per-machine last-successful-sync figure on the dashboard, and a local warning in the status line when this machine has not uploaded in N hours. The status-line half is the one that reaches the operator without them going to look. **Traps:** - `{sent: 0}` means four indistinguishable things (nothing to send / wrong host / wrong secret / lock held). Record success where a 2xx with a real `{accepted, rejected}` body was actually seen — NOT where `syncTail` returned. - A 200 alone is not success: posting to the app's base URL instead of `/api/sync` used to hit the dashboard page, return 200, and drop every event (R31). - Do NOT make the status line block on a network call. It renders every turn. - `/api/sync` is the only place that knows a sync succeeded, so server-side `lastSyncAt` keyed by machineId is the likely home — but the status-line warning may want a local source instead. That choice is the first decision. # F2 — Per-session turn timeline: carry cost vs new work The dashboard answers "how much" and, since #144, "which sessions". It does not answer **why a session was expensive**. Claude Code re-reads the whole conversation every turn, so most of a long session's cost is *carry* — paying again for context that already existed. On a 438K-token session that was 78% of each turn. The local status line already computes carry; the dashboard shows none of it, so the one number that would change behaviour (compact earlier, split the task, delegate) is invisible where spend is actually reviewed. **AC:** the existing session page (`src/app/session/[id]/page.tsx`) shows cost across the session's turns, split carry vs new. The useful shape is where the curve turns. **Traps:** - Check whether per-turn carry is derivable from already-stored fields (`inputTokens`, `cacheReadTokens`, `cacheWrite5m/1h`) or needs a new one. A new field must clear BOTH allowlists (mapper + `/api/sync`) and stay an aggregate — no transcript content, ever. - `src/app/dashboard-format.ts` compiles into the CLIENT bundle and must never import `src/parser/pricing.mjs`. - Sessions get long; `planDayBuckets` caps at 20 buckets for a reason. # F3 — Card titles are not headings Below the single `<h1>`, none of the six dashboard cards is a heading — shadcn's `CardTitle` renders a plain `<div>`. A screen-reader user has no heading structure to jump between. Surfaced by the e2e work (#142), recorded in the test-runner spec under "Known app-level gaps this surfaced". **AC:** real heading elements for card titles across dashboard, drill-down and session pages, asserted in e2e so it cannot regress. **Traps:** - base-ui's `Button` with `nativeButton={false}` stamps `role="button"` on the pager's anchors, so they are NOT `link`s in the accessibility tree despite being `<a href>`. Any a11y assertion must account for it. - No DOM harness in vitest (jsdom deliberately rejected on #132 — `getBoundingClientRect` returns zeros, producing false greens). Playwright is the only place this can be asserted. **Related bug in the same spec section:** selecting a date window then *immediately* clicking a split tab loses the window — a real race between two URL-writing controls. Low severity. Fold in or split out, operator's call. # Context Reasoning, findings and the traps above are in `claude-usage/tmp/split-drilldown/JOURNAL.md` while that workspace survives, and in `docs/features/*/spec.md` permanently.
claude-usage·main2 months ago