bytelyst-devops-tools

Author	SHA1	Message	Date
saravanakumardb1	df65b7a245	feat(agent-queue): report testing + optional autoship to the fleet (close testing->shipped) Previously the factory reported up to `review` and "shipping is always manual", so a coordinator job never reached a terminal stage autonomously. - On a passing local verify, always report `testing` to the coordinator so its stage reflects that QA passed (was stuck at `review`). - New AQ_FLEET_AUTOSHIP=1: the factory's verify gate IS the test phase, so advance the coordinator job testing -> shipped and land it in shipped/ locally. This closes the testing->shipped gap for an autonomous submit -> shipped pipeline. Default off keeps the human review gate authoritative (job rests at testing). selftest: +2 cases (autoship reports testing+shipped + lands in shipped/; autoship OFF reports testing but withholds shipped). Full self-test PASS. Generated with [Devin](https://cli.devin.ai/docs) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>	2026-05-31 04:21:44 -07:00
saravanakumardb1	8085501506	feat(agent-queue): extract Devin token usage from the conversation export Devin does not surface token/cost in its stdout or local log, so parse_usage previously emitted nothing for the devin engine (runs showed no metrics). Devin DOES expose per-step usage in its ATIF conversation export. - build_agent_cmd: pass `--export <path>` for the devin engine (path derived from the job log path so parse_usage can find it; harmless 4th arg for other engines). - parse_usage devin: read the export and sum per-step metadata.metrics input_tokens / output_tokens / cache_read_tokens; take model from agent.model_name. Pure grep/awk, no new dependency. USD cost is left unset (the export carries token counts but not cost) — the dashboard shows tokens + model, cost stays blank. These feed fleet_report_insights, so live devin fleet runs now report tokens + model to the coordinator (verified live: model "Claude Opus 4.8", tokensIn/out + cache populated on a real run). selftest: +1 case (parse_usage devin sums per-step tokens + model from --export). Full self-test PASS. Generated with [Devin](https://cli.devin.ai/docs) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>	2026-05-31 02:55:11 -07:00
saravanakumardb1	57831e3e7a	feat(agent-queue): report run insights to the fleet + normalize API base #1 fleet_report_insights: on a successful fleet run the factory now reports the parsed cost/token/effort metrics (model, tokensIn/Out/cached, costUsd, turns, toolCalls) plus the run result onto the coordinator run via POST .../lease/release (which also frees the lease). parse_usage already extracted these into the job meta; they were never sent. Engines that do not expose usage locally (devin) still land result + endedAt. #2 normalize AQ_FLEET_API: platform-service mounts fleet under /api, so a base without it silently returned 404 on every call. Strip a trailing slash and append /api unless already present, so AQ_FLEET_API=http://host:4003 works too. selftest: +2 cases (insights reported via lease/release; API-base normalization). Full self-test PASS. Generated with [Devin](https://cli.devin.ai/docs) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>	2026-05-31 02:27:51 -07:00
saravanakumardb1	dcf017a0de	docs(agent-queue): add run policy (isolated worktrees, least-privilege) Document how the daemon + agents must run after a review found jobs executing in --yolo/dangerous mode directly against live working trees (the root cause of repo dirtiness + duplicate commits). Policy: per-job worktree off origin/main, branch-per-task + PR, yolo:false by default (dangerous only in disposable sandboxes), clean-tree contract, one writer per repo. Linked from the README. Generated with [Devin](https://cli.devin.ai/docs) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>	2026-05-30 23:47:46 -07:00
Saravanakumar D	237481247e	docs(gigafactory): uppercase GIGAFACTORY folder + add index README Rename agent-queue/docs/gigafactory/ to docs/GIGAFACTORY/ and update every reference (README, system-overview code-map, and all phase job specs). Add an index README that lists the docs and points to the companion docs in learning_ai_common_plat. Docs-only; no behavior change. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>	2026-05-30 21:21:31 -07:00
Saravanakumar D	257efcb4bc	docs(gigafactory): consolidate gigafactory docs into docs/gigafactory/ Move GIGAFACTORY_ROADMAP.md and GIGAFACTORY_SYSTEM_OVERVIEW.md under agent-queue/docs/gigafactory/ so the scattered top-level docs are easy to discover. Update the README links, the overview code-map, and all phase job-spec source-of-truth paths to the new location. Pure docs move; no behavior change. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>	2026-05-30 21:01:23 -07:00
saravanakumardb1	1bcea394f5	chore(agent-queue): gitignore transient queue runtime state Jobs move through .state/inbox/building/testing/review/failed/shipped/logs at runtime, which constantly dirtied the repo and blocked clean rebases. Ignore the per-job lifecycle files (keeping each dir via .gitkeep) and stop tracking the consumed inbox job instances. Generated with [Devin](https://cli.devin.ai/docs) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>	2026-05-30 20:29:49 -07:00
Saravanakumar D	71e5ad6923	docs(gigafactory): add system overview with architecture diagrams; sync roadmap status Add GIGAFACTORY_SYSTEM_OVERVIEW.md — a current-state companion to the roadmap spec covering: what the Agent Gigafactory is, a completion snapshot, three Mermaid diagrams (component architecture, job-lifecycle state machine, atomic claim + lease-fencing sequence), the Cosmos data model, the scoring router, subsystem map, full /fleet REST surface, feature flags, the two control planes, a cross-repo code map, test coverage, next steps (Phase 4/5), and an honest bugs/gaps/risks section. All three Mermaid blocks validated with mermaid.parse. Also correct documentation drift in GIGAFACTORY_ROADMAP.md found during the review: - §0 progress table showed Phase 3 as "0% not started" while every Phase-3 box is ticked; updated phases 1-3 to done with realistic percentages. - Phase-2 boxes "scheduler/router wired into assignment", "tracker adapter direct call", and "factory enrollment + scoped tokens" are implemented in common-plat (coordinator.ts uses selectJob; routes.ts enforces enrollment.enforceFactoryToken; tracker-bridge.ts) but were left unticked — ticked with evidence and refreshed the stale "remaining for 100%" notes. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>	2026-05-30 20:11:02 -07:00
Saravanakumar D	66c91233da	feat(agent-queue): re-point TUI dashboard at /fleet API (parity) Add an opt-in fleet mode to the dashboard so an operator can drive the coordinator fleet from the same TUI used for the local folder queue. - lib/fleet-dash.mjs: dependency-injectable read/act adapter over the platform-service /fleet REST surface (jobs, metrics, factories, events, ship/requeue/reject). Pure-ish + fully unit-testable without a live service. - dashboard.mjs: render + act in fleet mode when AQ_FLEET_DASH=1 — board with counts, factories (per-factory rows or metrics aggregate), alerts, running (by lease/factory), actionable JOBS with manifest tags, recent, and a per-job events log. Single-flight async refresh keeps the last good board on failure; ship re-GETs a fresh leaseEpoch before PATCH; run/stop/promote are disabled (no safe server contract). Local mode is byte-for-byte unchanged. - lib/fleet-dash.test.mjs: 22 node:assert assertions (config, stage mapping, toBoard, fetch headers/timeout/errors, board assembly + graceful degradation, events, job actions) wired into selftest.sh. - docs: tick the Phase 3 "TUI re-pointed at /fleet" roadmap boxes. Verified: selftest.sh green (incl. new fleet-dash checks); live non-TTY render smoke against a stub /fleet server (both factories and metrics-aggregate paths); local mode unchanged. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>	2026-05-30 19:47:56 -07:00
Saravanakumar D	8a2270e0a6	feat(dashboard): surface manifest tags (priority/profile/caps/tracker) on the board Render a per-job tags line on the RUNNING workers and JOBS lists showing the routing inputs operators care about: priority, profile, capabilities, and the tracker-item reference. Tags come from the launched meta, falling back to the job's .md frontmatter for never-launched inbox jobs (new readManifest parser). The tracker-item becomes a clickable terminal hyperlink when AQ_TRACKER_WEB is set. Also renders the new budget_exceeded result as a failed RECENT row. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>	2026-05-30 19:27:41 -07:00
Saravanakumar D	7f77e9abc7	feat(agent-queue): enforce budget.wall as a hard wall-clock ceiling Parse the wall ceiling from the budget manifest map (budget: { wall: <dur> }) and arm it alongside the per-run timeout. Whichever ceiling fires first binds; the kill is recorded as result=timeout or result=budget_exceeded accordingly. budget.wall extends timeout: a job with only a budget.wall (no timeout) is now hard-killed at the ceiling. budget_exceeded is a terminal, non-retryable class by default and maps to the failed tracker status. Adds _budget_wall_secs + _effective_kill helpers (pure, unit-tested) and live selftest coverage; usd/tokens remain best-effort and are not enforced here. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>	2026-05-30 19:21:49 -07:00
Saravanakumar D	f1fe66fd4d	docs(roadmap): tick verified-done Phase 3 boxes (395-400,402) Phase 3 fleet control plane is implemented in learning_ai_common_plat: fleet API client, fleet map page, job table/detail/DAG/SSE/actions, cost burndown + multi-reviewer gate, scoring explainability, preemption, and Playwright fleet e2e. Box 401 (TUI re-point) remains open. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>	2026-05-30 19:13:25 -07:00
saravanakumardb1	a075a6ff30	Merge: Phase 2 two-factory parallel demo — exit criteria (§14) (#demo)	2026-05-30 01:58:55 -07:00
saravanakumardb1	0cde7def6a	feat(agent-queue): two-factory parallel demo — Phase 2 exit criteria (§14) Close the final Phase-2 exit-criteria box: >=2 factories executing jobs in parallel through one coordinator, proving the concurrency guarantees end-to-end. This is a DEMO HARNESS over the existing runtime — agent-queue.sh and lib/fleet-client.sh are unchanged (read + called, not modified). demo/two-factory-demo.sh: starts two real `agent-queue.sh run` daemons (mac-1 + ubuntu-1, separate queues/cwds) that compete ONLY through the coordinator, then asserts: (a) no double-assign — each of 3 jobs executed by exactly one factory; (b) fencing + reclaim — kill a factory mid-job, the reaper returns its job, the survivor reclaims + completes it, and the dead worker's late/zombie report (stale leaseEpoch) is FENCED (HTTP 409, never shipped); (c) parallelism — both factories hold active jobs concurrently. Dual-mode: CI-safe stateful stub by default; live platform-service when AQ_FLEET_API/AQ_FLEET_TOKEN set. demo/coordinator-stub.sh: stateful, mkdir-lock-guarded, file-backed coordinator implementing claim/lease/fence/renew/release + reaper-reclaim via the existing AQ_FLEET_API_CMD seam — the selftest stub pattern extended with shared state so >=2 processes coordinate through one coordinator. demo/README.md: stub + real invocations, env knobs, what each guarantee proves, what-to-watch guide. selftest.sh: +3 headless stub-mode checks (existing 68 unchanged byte-for-byte -> 71 total green). docs/GIGAFACTORY_ROADMAP.md: tick the §14 two-factory-demo box; annotate Phase-2 exit criteria; bump §0 Phase 2 to 80% (remaining: scheduler-core wiring [common-plat PR #31], tracker-direct call, factory enrollment). bash 3.2 + awk/sed/grep/pgrep only; mac+linux safe; no new runtime deps. Generated with [Devin](https://cli.devin.ai/docs) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>	2026-05-30 01:53:36 -07:00
saravanakumardb1	2d76af916d	docs(agent-queue): add Phase 3 overnight (10h) job — tunable scoring+preemption, DAG, budgets, tracker-web control plane	2026-05-30 01:48:39 -07:00
saravanakumardb1	08d8d715a1	docs(agent-queue): add Dependabot dependency-triage prompt for common-plat	2026-05-30 00:56:55 -07:00
saravanakumardb1	24fe1567f6	docs(agent-queue): draft Phase 2 next prompts — direct tracker->module wiring (§10) + two-factory parallel demo (exit criteria)	2026-05-30 00:40:21 -07:00
saravanakumardb1	fbecbe82b6	feat(agent-queue): fleet feature flags + shadow/dual-run (Phase 2) Add a safe, reversible path to validate the fleet coordinator against the proven single-host path BEFORE cutover, via three independently-toggleable flags: AQ_FLEET=0 pure offline (zero coordinator calls; offline path unchanged) AQ_FLEET_ROUTE=1 route_via_service: coordinator authoritative for claim (default = P2-S3) AQ_FLEET_ROUTE=0 local inbox authoritative (coordinator not used to source work) AQ_FLEET_SHADOW=1 dual-run (needs AQ_FLEET=1 + ROUTE=0): query coordinator in parallel, record divergence, NEVER act on it Precedence: SHADOW only when ROUTE=0; if ROUTE=1 + SHADOW=1, ROUTE wins (one-shot warning). lib/fleet-client.sh: fleet_route_enabled / fleet_shadow_enabled / fleet_flags_warn_once / fleet_flags_state; fleet_shadow_claim (read-only — isolated `-shadow` factoryId + dryRun, releases any real lease, never materializes), fleet_shadow_compare (AGREE/DIVERGE/COORD_EMPTY/LOCAL_EMPTY → .state/fleet-shadow.log), fleet_shadow_report (shadow:true, response never acted on), cmd_fleet_shadow_report (counts + agreement rate). agent-queue.sh: ROUTE-gate claim sourcing (claim only when route_via_service); shadow hook after the local authoritative decision each iteration (best-effort, error-swallowed — shadow can never fail a real job); `fleet-shadow-report` subcommand + help; resolved flags surfaced in `status`/`fleet-status`. tryClaim/fence/offline paths unchanged. Strictly side-effect-free on real job state: shadow never ships, quarantines, or mutates real jobs. Offline path byte-for-byte unchanged when AQ_FLEET=0. selftest.sh: +8 checks (shadow AGREE/DIVERGE/COORD_EMPTY, non-fatal 5xx, ROUTE precedence, ROUTE=0 local-authoritative, fleet-shadow-report summary, shadow_report unit). 60 prior checks unchanged → 68 total green. README + GIGAFACTORY_ROADMAP document the flag model + cutover ladder. Generated with [Devin](https://cli.devin.ai/docs) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>	2026-05-30 00:22:48 -07:00
saravanakumardb1	5c0ae020c0	docs(agent-queue): draft P2 prompts — factory enrollment+tokens (§12) + feature flags/shadow-dualrun	2026-05-29 23:52:14 -07:00
saravanakumardb1	21ebf8b1b7	docs(agent-queue): fleet integration section + roadmap P2-S3 ticks README: "Fleet integration (Phase 2)" — AQ_FLEET flag, env table, claim/heartbeat/ report/fence/renew protocol, offline-degrade + quarantine, offline-vs-fleet explainer. Roadmap: tick the Phase-2 §14 factory-agent item, add a P2-S3 slice note, bump §0 Phase 2 -> 55%.	2026-05-29 22:45:44 -07:00
saravanakumardb1	064dbf3d8f	test(agent-queue): fleet integration selftest cases (P2-S3) Adds 7 stub-driven fleet cases (AQ_FLEET_API_CMD stub, no live coordinator); never weakens the prior 53 (full suite now 60 green): - flag OFF (default): zero coordinator calls; offline job completes unchanged - register(heartbeat)+claim -> coordinator job materialized + executed to review/ - report+checkpoint: PATCH carries stage+leaseEpoch (+ wipBranch on building) - FENCING: stale-epoch 409 -> self-abort + quarantine (never shipped) - lease renew (unit): POST .../lease/renew with current leaseEpoch - offline-degrade: coordinator 5xx -> job completes locally (degraded), not quarantined - no-leak: bodyMd/token never appear in report payloads	2026-05-29 22:45:44 -07:00
saravanakumardb1	1d84712b47	feat(agent-queue): wire runner to fleet coordinator at minimal hook points (P2-S3) Sources lib/fleet-client.sh and adds a few fleet_enabled-gated hooks so the offline git-queue path is byte-for-byte unchanged when AQ_FLEET is unset/0: - cmd_run: register at loop start; per-iteration heartbeat (cadence) + lease renew for in-flight fleet jobs + claim one coordinator job into inbox when capacity. - meta: persist fleet_job_id + fleet_lease_epoch (from claim frontmatter). - run_worker: report `building` (with WIP checkpoint) after WIP setup and `review` before accepting the agent's output — a FENCED (stale-epoch/409) report self-aborts and quarantines (never ships); 5xx/unreachable degrades (finish locally). - _auto_echo: for fleet jobs route the outcome echo through the coordinator (fleet_events) instead of the direct tracker echo; offline jobs unchanged. - cmd_ship: fence-check before shipping a fleet job; release lease after. - status: show factory id + per-job fleet=<id>@e<epoch>; insights lists fleet_* fields. - dispatch + help: `fleet-status` command + a FLEET env section.	2026-05-29 22:45:44 -07:00
saravanakumardb1	a10d4003e6	feat(agent-queue): fleet coordinator client library (lib/fleet-client.sh, P2-S3) New sourced library implementing the factory side of the Phase-2 `fleet` coordinator contract — curl-only + POSIX awk, reusing the Slice-4 HTTP/JSON helper patterns, no new deps. Every function is a no-op unless AQ_FLEET=1. - fleet_enabled / fleet_api (AQ_FLEET_API_CMD test seam) / _fleet_call - fleet_detect_caps (reuses detect_capabilities) -> JSON caps array - fleet_heartbeat (+ _maybe cadence): registration == first heartbeat - fleet_claim: POST /fleet/claim, parse job id/bodyMd/leaseEpoch, materialize a transient local .md (fleet-job-id + fleet-lease-epoch in frontmatter) - fleet_report: PATCH fenced stage transition {stage, leaseEpoch, checkpoint?}; returns ok / FENCED(2, stale epoch -> self-abort) / degraded(1, unreachable) - fleet_lease_renew / fleet_lease_release / fleet_renew_active (fenced) - fleet_quarantine: park a reclaimed (fenced) job in failed/ for human triage - cmd_fleet_status: register + print factory identity/caps Report payloads carry only stage/epoch/checkpoint — never prompt/bodyMd/token.	2026-05-29 22:45:44 -07:00
saravanakumardb1	10395983e7	docs(agent-queue): draft parallel P2 prompts — scheduler/router core (§7) + fleet artifacts blob wiring (§13)	2026-05-29 22:32:41 -07:00
saravanakumardb1	9a073ef225	docs(agent-queue): draft P2-S3 factory-agent integration prompt (claim/heartbeat/report/fence behind AQ_FLEET)	2026-05-29 22:03:12 -07:00
saravanakumardb1	8ae504ca30	docs(agent-queue): tracker integration + close Phase 1 §10/§14 adapter (P1-S4) README: Tracker integration section (from-tracker/to-tracker, env config, label->manifest table, one-way-echo rule, AQ_TRACKER_AUTO, real-use note). Roadmap: tick §10 Phase-1 adapter items + the §14 tracker-adapter item; add P1-S4 slice note; §0 Phase 1 -> 95% (remaining: budget.wall + Node dash surfacing).	2026-05-29 21:35:16 -07:00
saravanakumardb1	1e0a17bbc0	test(agent-queue): tracker adapter selftest cases (P1-S4) Adds (never weakens) 7 stub-driven cases (AQ_TRACKER_API_CMD stub, no live service): from-tracker create + label mapping + idempotent; to-tracker shipped echo (PATCH done + metrics comment, asserts NO prompt body sent) + idempotent; HTTP 500 non-fatal; AQ_TRACKER_AUTO auto-echo on run. Full suite green (53 checks).	2026-05-29 21:35:16 -07:00
saravanakumardb1	b7a9ea1b7a	feat(agent-queue): tracker adapter — task <-> job round-trip (P1-S4) Implements §10 single-host tracker integration, closing the last Phase-1 §14 item: - tracker_api: one curl-only HTTP wrapper (base URL + bearer + productId header), overridable via AQ_TRACKER_API_CMD so tests need no live service. Emits the response body + a trailing HTTP-code line; _api_call splits into API_BODY/API_CODE. - aq from-tracker <ITEM_ID>: GET the Item, map title/description -> job body, labels (engine-class:/profile:/priority:/cap:) + Item priority -> frontmatter, and stamp tracker-item + a stable idempotency-key tracker-<id>. Materializes a .md into inbox/ via cmd_add; idempotent (Slice 1 dedupe) so a re-pull never dups. JSON parsed with POSIX awk (no jq) — mac + linux safe. - aq to-tracker <job>: one-way echo (child -> tracker, §24.5). PATCHes the Item status (building/review/testing->in_progress, shipped->done, failures->wont_fix, all overridable) and posts a metrics-only comment (result/attempts/duration/ tokens/cost/diff — NEVER prompt content or secrets). Idempotent via meta tracker_echoed; an echo failure (e.g. HTTP 500) is logged and non-fatal — the tracker is downstream, never authoritative for execution. - Opt-in auto-echo (AQ_TRACKER_AUTO=1, default OFF): the worker echoes on each transition (building via cmd_run, review/testing/failed via run_worker, shipped via ship/promote); never blocks or fails a job. - status + insights surface tracker-item and the last echoed status. curl-only HTTP; no new runtime deps; conventional + backward-compatible.	2026-05-29 21:35:06 -07:00
saravanakumardb1	d0348f23de	docs(agent-queue): P0 atomic-claim resolved (PR #29 ) — tick §4/§13/§14 fleet items	2026-05-29 21:05:38 -07:00
saravanakumardb1	2e9bd4dd1e	docs(agent-queue): record P2 Foundation merged + track P0 atomic-claim hardening (§4) - §4: implementation-status note — fleet module merged (PR #28); atomic claim NOT yet concurrency-safe (rev-CAS over unconditional write, sequential-only test) - add phase2-atomic-claim-hardening.md: updateIfMatch in @bytelyst/datastore (Cosmos If-Match + process-atomic memory) + concurrent claim tests	2026-05-29 20:43:28 -07:00
saravanakumardb1	0e94705ab7	docs(agent-queue): draft Phase 2 Foundation long-run prompt (fleet module + coordinator: claim/lease/fencing/reaper)	2026-05-29 19:54:33 -07:00
saravanakumardb1	e183919c60	docs(agent-queue): profiles + deps docs; tick §5/§6 + bump Phase 1 to 80% (P1-S2) README: Profiles & deps section (resolution precedence, persona, allowed-scope warn-only, deps/blocked + cycle detection); manifest table moves profile/deps/deps-mode to active. Roadmap: tick §6 catalog/persona/inheritance/allowed-scope and §5 deps + the §14 profile/deps/scope boxes; add P1-S2 slice note; §0 Phase 1 -> 80%.	2026-05-29 19:26:33 -07:00
saravanakumardb1	71d8a7cd4e	test(agent-queue): profiles + deps/DAG selftest cases (P1-S2) Adds (never weakens) temp-catalog + temp-git cases: profile verify inheritance + job-override precedence, persona-injection golden, profile capability inheritance, allowed-scope warn-only + path_in_scope unit, deps block->run, deps-mode soft (testing/), and submit-time cycle rejection. Full suite green (46 checks).	2026-05-29 19:26:26 -07:00
saravanakumardb1	f2dabdeb81	feat(agent-queue): starter profile catalog (P1-S2) profiles/<name>.md presets (name, persona, capabilities, default-verify, engine-class, prefers-engine, allowed-scope, review-policy) for developer, backend-engineer, frontend-engineer, ux-designer, ui-designer, qa, reviewer, docs-writer, and a reserved planner.	2026-05-29 19:26:26 -07:00
saravanakumardb1	3d99f04427	feat(agent-queue): profiles (persona + presets) and single-host deps/DAG (P1-S2) Implements roadmap §6 (profiles) and §5 deps on the bash runner, backward-compatible (jobs without profile/deps behave exactly as before). Profiles (§6): - profile_get / profile_persona / fm_eff helpers + PROFILES_DIR (AGENT_QUEUE_PROFILES override). A job's `profile:` inherits verify (<- default-verify), capabilities, engine-class, prefers-engine, allowed-scope, review-policy when the job omits them; job fields always override (precedence job > profile > default). Resolution runs via fm_eff inside the capability gate and resolve_engine, so inherited caps/engine-class take effect before launch. - persona injection: the profile's persona block is prepended to the stripped body fed to the engine (job .md unchanged on disk; nothing secret logged). - allowed-scope guardrail (WARN-ONLY): scope_check logs a non-blocking WARNING + records scope_warning= for changed paths outside the globs; path_in_scope is a pure, unit-testable matcher (`dir/**` = subtree). deps / DAG, single host (§5): - deps reference other jobs by idempotency-key. dep_satisfied: shipped/ (hard) or shipped/+testing/ (deps-mode: soft). deps_unmet drives a block-with-reason skip in inbox selection (never launched/failed); cmd_status surfaces "blocked (waiting on <keys>)". deps_would_cycle rejects cyclic submits on `add`. - _drain_pending: `--once` drains past dep-blocked jobs (idle can't satisfy them) while still waiting on retry/recovery backoff timers. Meta now records effective (inherited) capabilities/engine-class/prefers-engine/ review-policy/allowed-scope so `status` reflects resolved config.	2026-05-29 19:26:16 -07:00
saravanakumardb1	7c4f5bc9b0	docs(agent-queue): draft Slice 4 (tracker adapter) + Phase 2 Slice 1 (fleet data model)	2026-05-29 19:11:09 -07:00
saravanakumardb1	0443590ce4	docs(agent-queue): update Slice 2 prompt — branch off main (Slice 1+3 merged)	2026-05-29 19:05:34 -07:00
saravanakumardb1	87a4bf591a	docs(agent-queue): Resilience + Insights docs; tick §11/§25/§26 single-host (P1-S3) README: Resilience + Insights sections, retry frontmatter active (manifest table), retries_exhausted/recovered result values, recover/insights commands, honest token caveat. Roadmap: tick fully-completed single-host boxes in §11/§25/§26 with annotations; bump §0 Phase 1 to 55%.	2026-05-29 18:43:38 -07:00
saravanakumardb1	f46dd38adb	test(agent-queue): resilience + insights selftest cases (P1-S3) Adds (never weakens) temp-git-repo + stub cases: orphan recovery (+idempotent), WIP checkpoint/numstat, non-git skip, WIP resume, retry on verify_failed and crash (incl. no-retry when class absent), parse_usage extraction, per-engine aggregate. Inbox-empty-safe counts; avoids the pipefail+grep -q SIGPIPE trap.	2026-05-29 18:43:30 -07:00
saravanakumardb1	679d8b72cd	feat(agent-queue): dashboard insights column for finished jobs (P1-S3) Read-only from meta: tokens or cost + attempts + line deltas + duration; recognizes the new retries_exhausted result. agent-queue.sh stays the source of truth.	2026-05-29 18:43:30 -07:00
saravanakumardb1	1758bc1ab1	feat(agent-queue): single-host crash recovery, WIP checkpoint/resume, retry + insights (P1-S3) Implements the single-host bash equivalents of roadmap §25 (durability/crash recovery) and §26 (execution insights), plus §11 retry/dead-letter stand-in. Resilience (A1-A4): - recover_orphans + `recover` command: building/ jobs with a dead worker (dead pid, pidstart reuse-guard) are moved back to inbox/ with attempts incremented, on `run` startup and each loop. Idempotent (folder location is the guard). - WIP checkpointing: for a git cwd, _wip_start creates/checks out aq/wip/<job> and _wip_checkpoint commits changes on every exit path via an EXIT/INT/TERM trap; never commits to main/current branch; non-git cwd skipped. RESUME: a relaunch whose aq/wip/<job> exists checks it out first (continue from checkpoint). wip_base persisted in a write-once sidecar. - retry policy (now functional): retry { max, backoff, on } requeues failures whose class (timeout\|verify_failed\|crash) is in `on`, honoring backoff via next_eligible (selection skips until eligible), up to max attempts; exhaustion -> failed/ result=retries_exhausted with the WIP branch + full log preserved. - state integrity: all meta writes stay append-only; attempts/next_eligible/wip_* are re-derivable; recovery is crash-safe. Insights (B1-B6): - per-run metrics into meta: duration_s, exit, result, attempts, and (git cwd) files_changed/lines_added/lines_deleted from numstat wip_base..HEAD. - parse_usage(engine, log) adapter: generic AQ_USAGE line + Claude/Codex token heuristics; Devin/Copilot TODO; usage_estimated flag; never fabricates numbers. - status insights sub-line; new `insights [job]` command (per-job metrics or a recent table + per-engine token/cost/success/duration rollup). - privacy: only metrics are recorded, never prompt content or secrets. Backward-compatible: legacy .md and non-git cwd behave exactly as before.	2026-05-29 18:43:21 -07:00
saravanakumardb1	bc0c0e263c	Merge PR #1 : Phase 1 Slice 1 — evolved manifest, priority, capabilities, engine-class, idempotency Reviewed against the diff (capability gate before launch, 3-pass idempotency, priority+age selection, engine-class resolution, timeout/flock launch). selftest 18/18.	2026-05-29 18:12:41 -07:00
saravanakumardb1	1f18f5d7a3	docs(agent-queue): add Phase 1 Slice 3 prompt (resilience & insights, single host)	2026-05-29 18:10:43 -07:00
saravanakumardb1	beb225162a	docs(agent-queue): add durability/crash-recovery (§25) + execution insights/token accounting (§26) - §13: fleet_jobs stores instruction bodyMd (durable md SoT) + checkpoint; fleet_runs carries token/cost/model/tool/diff metrics - §25: instructions durable in Cosmos md, WIP checkpoint branch aq/wip/<jobId>, orphan detection, resume-vs-restart, fencing, retry->dead_letter, crash taxonomy - §26: per-run token/cost/latency/tool insights, honest metered-vs-estimated capture, rollups, control-plane surfacing, secret redaction - feature-catalog rows for §25 and §26	2026-05-29 18:09:32 -07:00
saravanakumardb1	470b2ce8d0	docs(agent-queue): version Phase 1 slice prompts (slice1, slice2) Track the delegated agent task prompts under docs/jobs/ so the slice decomposition of the gigafactory roadmap is reproducible and reviewable.	2026-05-29 18:05:06 -07:00
saravanakumardb1	67d8aa5766	docs(agent-queue): add work hierarchy & composite delegation (roadmap/epic) New §24 + feature-catalog row: - two delegation modes: atomic (leaf bug/feature/task) vs composite (roadmap/epic) - introduce job kind (leaf\|composite); composite routes to a planner/orchestrator that fans out child leaf jobs as a DAG across factories/agents/profiles - parentId hierarchy + rollup semantics (status/budget/verify/phase-gates) + idempotent re-run (skip shipped children) - source-of-truth/sync discipline (one record referenced by many; one-way echo) - HYBRID decision recorded: model kind/parentId/rollup in the fleet layer now, keep shared tracker ITEM_TYPES unchanged (label kind:roadmap), promote to a first-class epic type later via additive migration once proven - phasing: leaf-only P1-P2; manual composite P3; auto-decomposition planner P3->P5	2026-05-29 18:02:10 -07:00
saravanakumardb1	a9c69b1dce	docs(agent-queue): manifest field table (active vs reserved) + tick Phase 1 Slice 1 (P1-S1) - README: new "Manifest fields (Gigafactory Phase 1)" table marking ACTIVE vs RESERVED, capability-grammar table, idempotency-key semantics, copilot engine mapping, COPILOT_BIN, and capability_mismatch/no_engine result values. - GIGAFACTORY_ROADMAP: tick only the fully-completed P1 boxes (frontmatter parsing, capability detect+match, priority, backward-compat, capability grammar, engine-class taxonomy, idempotency-key semantics, README/progress), annotate partials, and bump §0 Phase 1 to in-progress 35%.	2026-05-29 17:44:37 -07:00
saravanakumardb1	4600a41e5d	test(agent-queue): self-test cases for manifest/priority/capabilities/engine-class/idempotency (P1-S1) Adds (never weakens existing) cases, each in its own temp AGENT_QUEUE_ROOT using the no-op engine stub: - backward-compat: legacy engine/cwd/yolo-only .md still lands in review/. - priority: with --max 1, a critical job queued after a low job runs first (order-recording stub). - capability mismatch: has:definitely-not-installed -> failed/ result=capability_mismatch, asserting the agent was never launched. - engine-class: agentic-coder + no engine, DEVIN_BIN stubbed -> review/. - idempotency: same key+body twice -> 1 inbox file; same key+changed body in inbox -> superseded; same key+different body after drain -> rejected. Inbox counts use find (not a globbing ls) so set -e/pipefail tolerate an empty inbox.	2026-05-29 17:44:27 -07:00
saravanakumardb1	0be5b34123	feat(agent-queue): evolved manifest, priority, capabilities, engine-class, idempotency (P1-S1) Implements Gigafactory Phase 1 - Slice 1 in the bash runner (backward-compatible; a legacy engine/cwd/yolo-only .md behaves exactly as before): - Parse all new §5 manifest keys via fm_get with safe defaults; record them in <job>.meta and surface priority/profile/capabilities/tracker-item in `status`. Only priority, capabilities, engine-class and idempotency-key are functional this slice; the rest (profile, prefers, budget, deps, deps-mode, retry, review-policy, artifacts, tracker-item) are stored but inert. - priority ordering: inbox_sorted picks critical>high>medium>low, ties by oldest; per-lock serialization preserved. - capability grammar + match: detect_capabilities advertises os/engine/node/has tokens; caps_match honors key, key:value, key<op>version and os:any. A job whose declared capabilities the host cannot satisfy is moved to failed/ with result=capability_mismatch and the agent is never launched. - engine-class resolution: explicit engine wins; else engine-class picks the first available engine honoring prefers-engine (agentic-coder->devin,claude,codex; chat-coder->copilot). No available engine -> result=no_engine. Adds copilot to the engine driver + COPILOT_BIN. - idempotency-key dedupe on add: same key+body -> no-op; same key+different body supersedes an inbox prior, else is rejected with a clear error. No change to queue/ data or the run/ship lifecycle. macOS + Linux safe.	2026-05-29 17:44:19 -07:00
saravanakumardb1	3ad9500623	docs(agent-queue): harden gigafactory roadmap after principal review Fix correctness/distributed-systems bugs and fill gaps in place: - atomic claim (optimistic concurrency/_etag), fencing token (leaseEpoch), coordinator-authoritative time added to core contract + scheduler + factory - lease reclaim via coordinator reaper, not Cosmos TTL (TTL only GCs rows) - split-brain/partition safety: fencing + distributed lock + quarantine - budget: wall is the only hard ceiling; usd/tokens best-effort (provider metering) - SSE live logs cannot use the buffering tracker proxy; use a streaming route + blob log storage (fleet_artifacts container) - manifest: capability grammar, engine-class enum, idempotency 409 + deps-satisfied semantics, dep cycle detection - tracker status mapping table + PR-flow ship semantics (merged+green vs pr-opened) - station/seat capacity, factory health definition, enrollment/bootstrap auth - Cosmos RU/indexing + claim-loop poll cost; add new sections: rollout/rollback & data migration (§21), capacity planning & cost (§22), ownership & RACI (§23) - success metrics now carry provisional SLO targets; Phase 2 checklist + index synced	2026-05-29 17:15:28 -07:00

1 2

67 Commits