f1-test-drive
Orchestrate F1 test drives to validate the Cyrus agent system end-to-end across issue-tracker, EdgeWorker, and activity rendering.
By cyrusagents · 616 installs
npx skills add cyrusagents/cyrus --skill f1-test-drive
Source repository · Upstream listing
F1 Test Drive
Run comprehensive F1 test drives that validate the full pipeline:
Issue tracker behavior
EdgeWorker execution flow
Activity rendering/output quality
Mission
Execute test drives that verify:
1. Issue tracker correctness
2. EdgeWorker worktree/session behavior
3. Activity output visibility and formatting
Test Drive Protocol
Phase 1: Setup
1. Create a fresh test repository (if needed):
2. Start F1 server:
3. Verify server health:
Phase 2: Issue Tracker Verification
1. Create test issue:
2. Verify issue ID and issue creation response.
Phase 3: EdgeWorker Verification
1. Start agent session:
2. Monitor activities:
3. Verify:
session started
activities appear
agent is processing issue
Phase 3.5: Slack Chat Session Verification (optional)
Use when validating the Slack → ChatSessionHandler → ClaudeRunner path. F1 exposes a test only endpoint /cli/dispatch chat that injects a synthetic app mention event without going through Slack signature verification ( SlackChatAdapter no ops Slack API calls when slackBotToken is undefined).
1. Dispatch a synthetic chat event:
The response contains a threadKey of the form <channel :<ts . Reuse the same thread ts to address the same chat thread on subsequent dispatches.
2. Verify shared auto memory wiring:
The chat workspace exists at <cyrusHome /slack workspaces/<sanitized threadKey / .
The shared auto memory directory exists (or is lazily creatable) at <cyrusHome /slack memory/ .
The claude query options event emitted by ClaudeRunner carries cqo.settingsAutoMemoryDirectory=<cyrusHome /slack memory .
3. Verify per thread workspace isolation alongside shared memory:
Dispatch a second event in a different channel/thread.
Confirm a separate slack workspaces/<other thread key / directory exists (workspaces remain isolated).
Confirm both dispatches' telemetry resolve to the same slack memory path (memory is shared).
Phase 4: Renderer Verification
1. Validate activity payload quality:
expected types (for example thought , action , response )
timestamps present
content well formed and readable
2. Validate pagination behavior:
Phase 5: Cleanup
1. Stop active session:
2. Stop background server process.
Reporting Format
Write report under apps/f1/test drives/ :
Pass/Fail Criteria
Pass when:
1. Server starts
2. Issue created successfully
3. Session starts and activities appear
4. Activity payloads are coherent
5. Session stops cleanly
6. No unhandled errors
Fail when:
server startup fails
issue creation fails
session does not start
no activities after reasonable wait
malformed activity data
unhandled exceptions
Important Notes
Prefer fixed port 3600 unless already in use.
Use fresh test repos per drive.
Preserve failed state when debugging.
For major runner/harness changes, run at least one F1 end to end validation before merge.
Multi Harness Note
This skill is intentionally harness agnostic:
Claude subagents can call this skill.
Codex/OpenCode workflows can reference the same skill content.
Harness specific adapters should be thin wrappers around this canonical skill.