SeekAlgo Axiom observability spec
**Approved design snapshot · Revised 30 September 2026**
SeekAlgo Axiom observability spec
Configuration update (2026-09-30): the current runbook supersedes the original per-worker credentials and opt-in flags. Vercel needs only its gateway URL and Next app token; downloader-owned services share one separate credential. Activation and release IDs are automatic.
Approved design snapshot · Revised 30 September 2026
Implementation follows this design in the linked observability PRs. The proposal language below records the original research; current verification and activation status are in runbook.md and the downloader companion PR. Production activation remains separate.
This proposal adds selected operational logs to SeekAlgo’s Next.js app and data downloader, including the browser backtest engines and the separate paper-trading runner. Its purpose is to explain which operation failed, at which stage, for which data or dependency, and whether it recovered. It deliberately avoids exporting the existing console streams.
Recommendation: one production Axiom dataset, small structured failure events, summaries from existing workflows, and a single exporter on the existing Hetzner server. The exporter aggregates and batches useful events. The design targets the Axiom free tier through at least 1,000 product users, with conservative daily-active-user scenarios. Logging continues as usage grows, and paid capacity is an acceptable growth path. The application keeps working when logging fails.
The first four sections explain the decision. The later sections define the events, safeguards, and verification required for implementation. All new behavior below is proposed, not implemented.
1 Proposed decisions
| Decision | Recommended choice |
|---|---|
| Scope | Operational logs only. No traces, metrics ingestion, Web Vitals, session recordings, prompts, scripts, market-data rows, or account/trade payloads. Existing PostHog analytics stays separate. |
| Axiom account | Personal free plan, initially one owner. Use one new production dataset named seekalgo-prod, subject to account entitlement and existing usage checks. |
| Retention | 30 days. Size retained data using the 1,000-user activity model and measured storage, without assuming a compression ratio. |
| Delivery | One authenticated collector and exporter container in the downloader deployment, independent of the Data API and application database. |
| Capacity goal | Stay within Axiom’s published free allowances through at least 1,000 product users by selecting useful events, aggregating routine work and measuring growth. No application-imposed daily/monthly log budget, export cutoff or reserved byte allocation. |
| Event policy | First meaningful failure, outcome or recovery transition, and compact aggregate summaries. Deterministic 1% samples of routine successful user operations. |
| Browser | Explicit parent-side operation failures through a restricted same-origin Next.js endpoint. No privileged token in the browser or isolated engine. |
| Alerts | Three aggregate monitors covering critical failures, stalled/data-dependent work, and exporter silence. Email to the account owner; no notifier is connected during this research. |
| Implementation gate | Review this spec first. Product code, deployment settings, datasets, tokens, alerts, and external integrations have not been changed. |
User direction: preserve useful observability as the product grows. Application log delivery must not stop when a forecast, usage warning or internal accounting counter reaches a value. Changes to paid capacity can accommodate growth; activity assumptions and real measurements determine when that becomes useful.
2 Research scope and evidence
The audit used clean snapshots of current main because the original local checkouts were behind recent changes and contained unrelated work.
| Repository | Audited commit |
|---|---|
| sa_next_app | 040a85c91f8f5c8c0f01b2d77bece105be8d8f07 |
| sa_data_downloader | 88f881e5f687f92086731ae39c7c179912c1be09 |
The review covered route and stream outcomes, builder/provider execution, browser engines, market-data loading, auth and persistence, MCP, downloader scheduling and backfills, Dhan streaming/recovery, Parquet, the Node paper runner, Compose ownership, and deployment entry points. Public Axiom and Next.js documentation was checked separately.
This is source evidence, not a live production incident report. Production release IDs, actual traffic, existing Axiom account usage, server resource headroom, and measured event sizes remain rollout checks. No production jobs, migrations, provider requests, or full data scans were run.
Findings that change the logging design
| Finding | Product implication and proposed observation | Source evidence |
|---|---|---|
| Builder failures can be returned inside an HTTP 200 stream; some unexpected exceptions become 400. | Record the terminal semantic outcome and failed stage. HTTP status alone will miss these failures. | Builder route, stream handling |
| Provider catches lose diagnostic detail; usage finalization can fail after the primary operation. | Preserve provider status/error category and primary versus secondary failure. Unknown billed usage must remain unknown. | Pi stage catch, AI finalization |
| Browser and sandbox execution do not run inside a Vercel request. | Observe the parent’s terminal engine result and result-save outcome separately. Preserve isolation. | Isolation, worker hook |
| Existing consoles can include generated code, user-script output, and bot account data. | An unrestricted drain would capture unwanted content and noise. Create events from an allowlist. | Chart generated code, runner account output |
| Downloader health checks do not prove usable data; a fresh runner heartbeat can coexist with a DB error. | Keep process liveness, dependency readiness, progress, and coverage as separate fields. | API health, runner health, bot discovery error |
| Admin backfill can break on a provider failure, write accumulated partial data, then report completed. | Emit requested versus checked bounds, termination reason, publish mode, and observed partial outcome alongside the current job status. | Break on failed result, overwrite and completion |
| Parquet publication and DB bookkeeping are separate stages. | Distinguish published data with pending bookkeeping from fetch or storage failure. | Scheduler publication, batch bookkeeping |
| Stream/recovery failures can remain in local state; queue rejection is easy to overlook. | Export compact state changes and backlog/rejection counts, rather than every contract or bot attempt. | Options heartbeat, queue rejection |
These are code-path concerns. Their production frequency has not been measured.
3 Runtime ownership
| Runtime or owner | Observability responsibility |
|---|---|
| Next.js server routes and MCP | Request outcomes, builder/provider stages, auth/dependency failures, database persistence, paper launch. |
| Browser application parent | Data load, worker/sandbox completion, result-save transport, fatal app/chart failures. Browser assertions remain unverified. |
| pipeline | Non-Dhan downloads and incremental scheduling. |
| nse-pipeline | Dhan 5-minute pipeline. |
| nse-slow-pipeline | Other Dhan timeframes. |
| nifty-options | Fixed-contract option streaming, authoritative recovery, and publication. |
| api | Data/symbol/live routes plus the in-process admin backfill manager. |
| nifty-equity-backfill | Finite historical equity work. A confirmed completion may legitimately stop this container. |
| nifty-option-archive | Daily EOD archive work, repairs, and scheduled waiting. |
| live-runner | Separate Node paper execution, queueing, engine workers, and reconciliation. |
| Proposed log-gateway | Collector, aggregation, deduplication, bounded queue, observed usage, Axiom acceptance, and transport health. |
| Proposed host lifecycle probe | Fixed allowlist of container states, exits, OOM and restart counts, including failure before an app logger starts. |
The eight existing containers are defined in Compose. The paper runner starts from Dockerfile.runner, not from the Next.js checkout. Its scheduled work belongs to this deployment; the Next.js bot run endpoint is a manual path.
Docker already rotates local logs at 50 MB × 5 files per service. Keep that local fallback and do not forward it wholesale. Failed Vercel builds, GitHub deploy jobs, image pulls, and host faults also have native logs; application event delivery does not replace those records.
4 Axiom constraints that matter
| Published Personal allowance | Design consequence |
|---|---|
| 500 GB/month data loading | Plenty of headroom for selected events; burst and request limits still apply. |
| 25 GB stored data | Storage and ingestion are separate. Storage is compressed; no fixed compression ratio is assumed. |
| 10 GB-hours/month query compute | Monitor runs and dashboard queries consume this allowance too. |
| 30-day maximum retention | Keep the proposed dataset at 30 days. |
| 3 datasets and 256 fields per dataset | One dataset, a fixed schema, and no dynamic field names. |
| 1 user and 3 monitors | One owner for this rollout. Team access would require revisiting the plan. |
| Email and Discord notifiers | Start with email only. |
Sources: pricing and plan limits. Exact Personal storage-overflow behavior is not specified in the reviewed documentation; this design does not rely on automatic eviction or a precise pause at the ceiling.
Axiom charges API ingestion against the uncompressed HTTP request. Query compute reflects execution time and allocated memory, rather than just bytes scanned. Compression does not reduce the ingestion allowance. See usage accounting.
Ingest HTTP 200 can still report failed events. Check ingested and failed counts, failure details, and processed bytes. A successful SDK flush is insufficient acceptance evidence. See the ingest API.
Delivery approaches considered
| Approach | Benefit | Tradeoff | Decision |
|---|---|---|---|
| Selected events through one exporter | Shared aggregation, persistent delivery queue, one token, central storm control and usage visibility; independent of the app DB. | One extra small container and a shared delivery dependency. Pre-collector events may be lost. | Recommended for consistent summaries and delivery. |
| Direct Axiom SDK in every runtime | Less infrastructure and quick initial wiring. | Distributed aggregation and transport visibility require more coordination; in-memory buffers are not durable. | Valid simpler alternative if central aggregation is unnecessary. |
| Automatic Vercel drain or blanket Docker forwarding | Broad runtime coverage with little instrumentation. | Includes unwanted console content, successful traffic and noisy loops; weak control over useful information. | Exclude from this rollout. |
The Vercel integration uses a log drain. The SDK guides provide integration mechanisms, not a requirement to capture every request or every logging record.
5 Proposed architecture
flowchart LR
B[Browser parent] --> R[Restricted Next.js relay]
N[Next.js server events] --> G[Independent log gateway]
R --> G
P[Python owners] --> G
J[Node paper runner] --> G
H[Host lifecycle metadata] --> G
G --> S[Validate and redact]
S --> D[Coalesce repeated failures]
D --> Q[Bounded persistent spool]
Q --> A[Axiom seekalgo-prod]
A --> U[Observe usage and forecast growth]
The gateway lives in the downloader deployment but has its own process, persistent state directory, resource limits, and ingest endpoint. It must not use the application PostgreSQL pool, Dhan recovery SQLite database, or Data API request handlers.
Inside Compose, producers send small batches over the internal network. Next.js uses authenticated HTTPS through a dedicated gateway route at the existing edge proxy. Only the gateway holds a dataset-scoped Axiom ingest token. Producer authentication uses separate service credentials; the gateway derives service identity from those credentials and applies a service-specific event allowlist.
Application delivery
- Next.js registers the supported after/waitUntil lifecycle hook in request context before work begins. It finalizes the event after stream processing and persistence settle; returning the streaming Response is not completion. Gateway delivery has a 2-second timeout and no producer-side network retry loop. It must not delay the response, mask the primary exception, or use the one-connection production DB pool.
- Python and Node workers use a small background queue, bounded at 128 events or 256 KiB per process. Flush every 10 seconds, on batch fullness, and best-effort shutdown. The collector owns Axiom retries.
- Browser events use a same-origin endpoint accepting at most 5 events and 10 KiB per call. The relay sets trust=client_unverified and derives service/actor metadata; it never accepts client-supplied authority. Keep client_release separate from relay_release so an old browser bundle is not relabeled as the current server release. Validate names, enums, IDs, numbers and origin. Authenticated intake must avoid getCurrentDbUser, which depends on the app DB.
- Anonymous intake is limited to app-fatal and auth-dependency failure categories, contains no user identity, and is marked client_unverified. Origin validation is an abuse check, not authentication. Per-source validation, duplicate suppression and abuse controls protect diagnostic quality and availability.
- Browser intake coalesces repeated copies of the same failure while retaining first occurrences, occurrence counts and recovery. Authentication, origin checks, bounded payloads and abuse protection apply per source; legitimate new outcomes are not limited by a global daily event allowance. Temporary IP data used for abuse control is not exported to Axiom.
Next.js lifecycle callbacks remain within the platform’s duration limits. A hard timeout, OOM, browser close, network loss, or outage before gateway acknowledgement can lose logs. The persistent spool protects only events the gateway has accepted. See Next.js after and Axiom Next.js guidance.
The host lifecycle probe runs outside the gateway container using existing host ownership. It reads only fixed container-state metadata and forwards through the same selected-event pipeline. Do not mount the privileged Docker socket into the gateway or ship raw startup logs. A boot failure produces an exit/restart event; its detailed traceback remains in native Docker logs.
6 Event policy and schema
Use a typed event constructor in each language. Build from allowed fields before serialization; do not serialize arbitrary Error/request objects and rely on regex cleanup afterward.
Every event includes _time, schema_version, event_name, event_id, service, component, environment, release, severity, outcome and trust. Set _time to the actual event time; delayed ingestion must not make old failures appear new. Derive alert_class (none, critical or degraded) in trusted code from the reviewed event taxonomy, not browser assertions.
Optional fields are selected by event family:
| Group | Allowed examples |
|---|---|
| Correlation | operation_id, request_id, parent_request_id, failure_id, session_id, revision_id, run_id, job_id, cycle_id, batch_id, data_load_id, intent_id. |
| Failure | static reason_code, error_code, safe exception class, dependency, provider_status, stage, retryable, attempt, fingerprint, recovered. |
| Timing | duration_ms, last_success_at, last_progress_at, closed_bar_time, oldest_pending_age_s, queue_age_s. |
| Data identity | venue, provider, market_type, data_type, timeframe, source_timeframe, scrip_id, canonical instrument/contract key, dataset_id, source_basis. |
| Counts | attempted, succeeded, failed, partial, deferred, empty_unverified, rows_published, retries, queue_rejected, suppressed, dropped. |
| Runtime | engine version/hash, client_release, relay_release, configured owner, process_alive, dependency_ready, coverage_state and coverage_basis. |
| Sampling | sampled, sample_rate, interval_start and interval_end for aggregate summaries. |
Keep the full registry below 128 fixed flattened fields across the dataset. Arbitrary nested dictionaries, dynamic contract-keyed objects, lists of all bots/contracts, and unknown fields are rejected. Per record: target below 1 KiB, maximum 2 KiB UTF-8 including its envelope. Diagnostic stack locations are optional, capped to 8 sanitized frames within that same size limit. Exceeding optional detail is truncated with a flag; an invalid core envelope is dropped and counted.
Example proposed event:
{
"_time": "2026-09-30T04:20:00Z",
"schema_version": 1,
"event_name": "runner.dependency.changed",
"event_id": "f4e81d5f-12d3-4f32-a327-cc5ae5e0cd74",
"service": "live-runner",
"component": "bot_discovery",
"environment": "production",
"release": "88f881e",
"severity": "error",
"outcome": "db_failed",
"trust": "server",
"reason_code": "db_connection_unavailable",
"dependency": "postgres",
"process_alive": true,
"dependency_ready": false,
"duration_ms": 5000,
"recovered": false
}
This is an illustrative record, not a captured incident. Log timestamps use UTC; human reports and usage forecasts use Asia/Kolkata. Account reset headers are interpreted independently.
What must stay out of Axiom
Prompts, conversation history, thinking/activity text, tool arguments/results, generated or user code, compiler source excerpts, script console output, OHLCV/ticks/order books, trade/equity arrays, balances, positions, P&L, quantity/price details, credentials including encrypted ciphertext, PIN/TOTP/client IDs, cookies/JWTs, database URLs, SQL and parameters, environment dumps, raw provider bodies, user email/name/IP, and arbitrary URLs/query strings.
Financial outcomes remain in their existing product records. Axiom is diagnostic history, not a trading or audit ledger. Safe token/cost counts may accompany AI operation outcomes only when known; failed usage is nullable.
7 Next app logging locations
P0 means first implementation priority; P1 means the next layer. Priority is distinct from event severity. Each row inherits the schema and privacy rules above.
| Priority and event family | Exact boundary | Information worth recording | Noise rule |
|---|---|---|---|
| P0 request.finished | API route terminal handling and existing route catches | Route template, action, status, semantic outcome, duration, failure_id; unexpected dependency errors even if current status is 400/401. | Unexpected failures and slow outliers; 1% success sample. Exclude assets, health, OPTIONS and telemetry intake. |
| P0 builder.message.finished | message persistence, builder stream | Accepted/saved/refused/clarification/cancelled/timed_out/failed/disconnected; failed stage, saved version, model and usage-known state. | One server terminal event. Include bounded stage/check results; no progress deltas. Ordinary clarification/refusal is informational. |
| P0 ai.stage.finished | Pi provider adapter, AI orchestration | Classification/authoring/repair/review, provider status, context limit and estimated bytes, duration, safe reason, nullable usage. | Usually fold stages into the terminal message; retain failed stages with shared failure_id. No SSE chunks. |
| P0 builder.check.finished | validation and review | Catalog/coverage/compile/review outcome, attempted repair, exhausted generated-code validation. | Successful repair is a recovered warning. User syntax errors and business denial are not platform outages. |
| P0 data.coverage.checked | market check, planning check | Available/missing/unverified/unsupported, canonical identity, status, timeout/decode/guard reason, warmup and requested bounds. | One operation summary with counts and at most 3 failed identities; infrastructure unverified differs from real 404. |
| P0 backtest.lifecycle | server run and save, parent execution | Accepted, engine_finished, result_saved or status_save_failed; run/revision IDs, duration and stage. | Retain accepted and terminal transitions. Successful computation followed by failed save is a persistence failure. No progress events. |
| P0 backtest.engine.finished | isolated runner, worker hook | Runtime/engine hash, bar and symbol counts, timeout/crash/cancel/invalid result/user-script error. | Exactly one parent outcome. Never add network access to the sandbox. Routine generic successes sampled. |
| P1 data.browser_load.finished | market data load, DuckDB init | Full/slice/static source, candidate statuses, bytes, load/decode stage, cache disposition and canonical dataset. | Owner load event only; sample successes, keep failures/slow loads. Aggregate followers/cache activity. |
| P1 browser.operation.failed and chart.render.degraded | chart orchestration, builder error boundary | App component, operation, Next digest, safe fingerprint, chart fallback/recovery, release. | Explicit app failures only; browser assertions unverified. One event for a fallback, not every effect/retry. |
| P1 db.operation.failed and ai.usage.finalized | pool, builder transactions, usage finalization | Operation/phase, SQLSTATE category, connection wait, rollback failure, primary outcome and secondary usage-write failure. | No per-query success events or DB-backed logger. Preserve cause even when the current catch replaces it. |
| P1 auth.operation.finished and credentials.operation.finished | auth, credentials route | Mechanism, action, upstream/config/DB-sync/crypto failure categories. | Aggregate ordinary expiry/invalid/denied states. Exclude all request/result credential data. |
| P0 mcp.tool.finished | tool wrapper, route bridge | Tool enum, per-tool operation ID, semantic isError, bridge status, duration, safe failure stage. | HTTP 200 can still fail. No tool args/results, raw bodies, or request-by-request MCP polling. |
| P0 paper.launch.finished | transaction and idempotency | Created/replayed/rejected/storage_failed; launch idempotency ID, run/revision/bot reference. | Success only after commit. Replay informational. No name/source/capital/settings. |
| P0 bot.tick.finished and paper.reconciliation.changed | manual tick, reconciliation | Trigger, skip reason, quote basis/age, processing and persistence result, intent transition. | No-new-bar/paused/already-processed are ordinary counts. Quote failure with pending work is actionable. Production scheduled summaries come from the Node runner. |
| P1 release.started and engine.compatibility | engine sync manifest, migration runner | Release, engine/source hash, migration outcome and safe tag, config-presence booleans. | One per release/process or compatibility change. Native build/deploy logs remain fallback for failures before runtime. |
Current main has no payment/Stripe flow or Vercel cron configuration to instrument. Do not invent billing or cloud cron events.
8 Downloader and paper runner logging locations
| Priority and event family | Exact boundary | Information worth recording | Noise rule |
|---|---|---|---|
| P0 download.batch.finished | scheduler publish, DB batch update | Cycle/batch IDs, fetched/published/bookkept counts, published_db_pending, stage duration and safe failure. | One summary/window; first failure and recovery. Avoid fetch plus handler duplicate stacks. |
| P0 backfill.job.finished | admin backfill termination, publication | Requested/checked/observed bounds, provider termination, complete/partial/failed/empty_checked, current job status, overwrite/append mode. | One terminal event, plus aggregated repeated retries. Do not label current completed status as verified coverage. |
| P0 service.readiness.changed | API health, runner health | Alive, dependencies ready, last successful progress, coverage basis, owner state. | Transitions plus compact periodic summary. Health polling itself is excluded. |
| P0 data_api.request.failed | Parquet serving, filtered reads, live data | Resolve/read/filter/encode/send phase, status, request ID, dataset, requested bounds. | Observe ASGI send completion, including late FileResponse failure. Expected missing/unavailable data and invalid input have distinct reasons. |
| P0 provider.auth.state.changed and provider.request.failed | Dhan auth, Dhan retry/cooldown | Safe operation, status, auth refresh state, expiry, retry count, cooldown. | First transition and recovery, then grouped counts. Never log auth URLs/body or response text. |
| P1 provider.budget.state.changed | shared provider budget | Deferred requests, live/background quota classes, cooldown and threatened backlog. | Controlled deferral informational; warning only when required progress is endangered. |
| P0 stream.connection.changed and options.recovery.summary | stream ownership, recovery and progress | Connections/subscriptions, last message/trade, urgent/recent/history counts, unresolved traded buckets, oldest repair, progress persistence failure. | No packets/ticks/per-contract successes. Keep state changes immediately; aggregate routine counts every 15 minutes. |
| P1 scheduler.state.changed and scheduler.window.summary | main loop, task selection | Empty config, disk pause, due/dispatched/failed/published/deferred, oldest due age, owner filters. | No warning every 3 or 10 seconds. One first fault, rollups and recovery. |
| P1 provider.data_validation.failed | result validation, Dhan timestamp normalization | Empty-window reason, raw/closed/filtered counts, schema/timestamp collision, source versus derived timeframe. | Group by error family, not every option contract. Unknown empty data is not assumed no-trade. |
| P1 storage.operation.failed and dataset.checkpoint.diverged | atomic writer and lock, checkpoint reads | Read/write/lock/replace phase, corruption/missing distinction, rows changed, file bytes, DB versus Parquet checkpoint. | First failure/change and write totals/window. No INFO event for every rewritten file. |
| P1 archive.partition.assessed and import.run.finished | EOD partition work, import worker, historical downloader | Checked/published/unavailable dates, duplicates/errors, partial/resumable cursor, archive basis and next scheduled work. | One partition/run summary and terminal transition. Scheduled/running and saved-month counts are not completeness. |
| P1 catalog.sync.finished | identity mapping, reused provider IDs | Added/expired/reused mappings, canonical contract and generation; available/missing/empty/corrupt counts. | One sync summary with bounded failed examples. No whole catalog dump. |
| P0 runner.execution.failed and runner.window.summary | bot processing, cycle | Discovery/data/closed-bar/alignment/quote/engine/persistence stages, due/skipped/processed counts and durations. | No per-bot successful tick log. Aggregate 15 minutes; first failure and recovery use stable bot references. |
| P0 runner.queue.overloaded and runner.engine.failed | queue rejection, engine pool | Rejected/retry/exhausted deltas, oldest age, worker timeout/crash/replacement reason. | Queue deduplication and normal worker recycling are expected; overload and timeout distinct. |
| P0 runner.reconciliation.changed | reconciliation, queue admission | Intent created/executed/expired/blocked, no quote, pending age and persistence outcome. | Transition and action counts; no account, desired-position or signal payloads. |
| P1 runner.startup.failed and worker.lifecycle.changed | runtime checks, startup/fatal loop, deployment | Engine checksum/schema compatibility, release, exit/OOM/restart metadata and declared worker completion. | Explicit safe error extraction; JSON.stringify(Error) can lose detail. No CI stdout or environment upload. |
Instrumentation reads existing workflow counters and outcomes. It adds no provider fetches, history repairs, full Parquet scans, archive audits, or new calendar construction in the logging path.
9 Correlation and data correctness
IDs across both systems
One operation_id follows a user action through builder requests, data loading, engine completion and save. Each HTTP attempt also gets a request_id. A downstream owned-service call has a new request_id and parent_request_id; pass operation_id separately. Carry run_id/revision_id across the later result-save request.
Use validated opaque IDs. Incoming browser IDs are hints, not authority. The gateway overrides service identity; authenticated servers derive actor_key as a keyed pseudonym when necessary. Do not use a module-global “current user/request” under concurrent Vercel requests.
The paper-launch requestId already acts as an idempotency key; name it launch_request_id in logs so it cannot be confused with HTTP request_id. Scheduler cycles get IDs before provider work, not only after a download job row exists. Shared runner data-cache fetches use data_load_id with consumer links; they must not borrow an arbitrary first bot’s execution identity.
For owned HTTP traffic, the caller generates a new downstream request UUID and sends it as X-SeekAlgo-Request-Id. The receiver validates and retains it, or generates a replacement for invalid/missing input. X-SeekAlgo-Parent-Request-Id carries the caller request UUID; X-SeekAlgo-Operation-Id carries the action UUID. Return the receiver's request ID in the response. Parent and downstream IDs must differ.
Direct browser-to-Data-API Parquet loads will incur CORS preflight with new headers. The existing API allows headers but does not expose a response ID. Validate Caddy, CORS, range/file responses and caching before enabling this path. Until then, browser dataset/time association is approximate, not exact end-to-end correlation.
Canonical market identity
Derive dataset_id from existing canonical fields: venue, instrument or exact contract, market_type, data_type, timeframe, and distinct source basis/version where applicable. Keep provider separate: NSE is the venue; Dhan is the provider.
Options include underlying, expiry, canonical strike and CE/PE. Provider security IDs can be reused and are not sufficient identity. Spot, perpetual/dated futures, EOD option archive, intraday candles, and rolling research datasets remain distinct. Preserve scrip_id/source_scrip_id and normalized relative storage keys when available.
The legacy browser cache key omits venue (source). Do not reuse it as the logging identity or silently change cache behavior in this implementation.
Coverage is evidence based
- Separate process_alive, dependency_ready, progress_recent and coverage_state.
- Classify known closed session, before listing, after expiry, confirmed no trade, forming bars filtered, provider failure and empty_unverified separately.
- A provider empty response does not establish no-trade; an archive 404 does not establish a holiday.
- For options, use observed traded buckets or authoritative traded-contract evidence when already available. Registered/subscribed contract counts are not candle completeness.
- Requested/checked bounds differ from published bounds. Cursor movement, max timestamp and cumulative rows are not proof of unique coverage or absence of interior gaps.
- State coverage_basis and checked_at. If no trustworthy session/calendar evidence is available, leave coverage unverified and avoid guessed missing-bar incident alerts.
- Current Dhan scheduling uses fixed weekday/session rules (source); logging must not present those as a full exchange holiday/special-session calendar.
- scripts/validate_dhan_nifty.py fetches and persists index data; it is not an all-strike option completeness audit and was not run.
No data-source substitutions, fabricated bars, domain-state corrections or new completeness guarantees are authorized by this spec.
10 Efficient logging and capacity for 1000 users
Selection and aggregation
Use one failure_id for a causal failure crossing layers. A detailed failure plus a parent outcome may both be useful, but repeat stack details only once.
The storm fingerprint uses event family, service/component, route/action, provider, stage and normalized reason/error category. Exclude user/job/bot/contract IDs and free-text messages. Otherwise thousands of contracts would each create a “first” error.
Export the first meaningful occurrence per fingerprint in a 10-minute window, then a rollup containing suppressed_count and bounded affected examples. Retain recovery transitions. Bound in-memory fingerprint cardinality for process reliability; overflow groups by service/error family and retains an overflow counter. This coalesces repetitions while preserving distinct actionable failure categories and their occurrence counts. Exact affected counts are emitted only if actually maintained; saturated counts are flagged as lower bounds. Exemplars are capped to 3.
Routine successful requests use a deterministic 1% event sample. Builder message outcomes, accepted/terminal backtests, paper launch and import transitions are retained. Successful AI stages are folded into the terminal message instead of producing multiple logs. Successful bot polls and per-bar/tick/progress messages are represented by summaries, not individual records. Sampling and event policy do not change automatically as account usage accumulates.
For Next.js, each eligible dynamic request forwards a compact counter delta in its completion-hook batch; the gateway combines those deltas into 15-minute route/outcome summaries across instances. It does not depend on an instance surviving for 15 minutes or on a background timer. Workers aggregate existing workflow counters locally. Successful counters are not exported as one Axiom event per request.
Gateway counters describe received observations; loss before collector acknowledgement or worker termination can make product denominators incomplete. Mark completeness=best_effort for Next.js/browser counts and dropped/lost intervals. Do not label these as exact product success rates.
Remove incoming duplicate event IDs before aggregation using a bounded 24-hour deduplication set. Individual causal failure investigations count one authoritative boundary and unique failure_id. Rollups have a stable rollup_id, nonoverlapping interval and delta counts of suppressed occurrences; first-occurrence exemplars are separate. Deduplicate rollup_id before summing deltas. A saturated deduplication set flags count quality as degraded.
Slow-operation initial thresholds: 2 seconds for ordinary server/data operations and 30 seconds for an AI/backtest operation. Longer expected historical jobs are judged by progress age, not total job duration. Tune these diagnostic thresholds from baseline measurements.
Workload assumptions
The product target is at least 1,000 users. Registered users do not all use the product every day, so the conservative capacity cases also size 1,000 daily active users. Product users are separate from Axiom’s one-owner free-plan entitlement.
Model one builder terminal record per message, three records per backtest (accepted, parent engine terminal and server result saved/failed), all other API failures, a 1% sample of other successful API requests, and 2,000 infrastructure/route/failure-rollup records daily. Assume 2% of other API requests fail. The 2,000 figure is a planning assumption that includes periodic summaries and selected failure detail, not a daily event allowance. Additional distinct failures may increase it.
For paper trading, assume one running bot per daily active user and three selected records per meaningful action/intent lifecycle. Count 10 such actions per bot/day in the ordinary cases and 20 in the heavy case. Also assume each bot evaluates every five minutes over 24 hours: 288 evaluations/day, with 2% failing and three diagnostic/outcome records per failed evaluation before any storm-coalescing savings. This is more conservative than an NSE-only session. Routine polling and successful per-bar execution remain aggregate counts. These assumptions size log volume; they do not establish that the current execution host can run 1,000 bots.
Estimated selected log volume
Assumptions are per active user/day; all other API request counts exclude the builder/backtest requests already modeled. GB below means decimal gigabytes, and KiB means 1,024 bytes. These are transparent planning estimates, not measured traffic.
| Case with 1000 product users | Daily active users | Messages / backtests / other requests | Selected events/day including paper actions | Payload over 30 days at 1–2 KiB/event | Planning payload with 2× margin at 2 KiB/event |
|---|---|---|---|---|---|
| Ordinary adoption | 200 | 10 / 2 / 100 | 15,252 | 0.47–0.94 GB | 1.87 GB |
| All users active daily | 1,000 | 10 / 2 / 100 | 68,260 | 2.10–4.19 GB | 8.39 GB |
| Heavy daily use | 1,000 | 50 / 10 / 100 | 162,260 | 4.98–9.97 GB | 19.94 GB |
Formula: daily_active_users × (messages + 3 × backtests + other_requests × (98% × 1% + 2%) + 3 × paper_actions + 288 × 2% × 3) + 2,000 selected summaries/details. Multiply by event bytes and 30 days. The 2× planning margin accommodates framing, retries and extra distinct diagnostic events; continuous retry storms or activity beyond these assumptions can exceed it.
The heavy 1,000-daily-active-user case plans about 19.9 GB/month, approximately 4.0% of the published 500 GB ingestion allowance. With 30-day retention, the same steady workload introduces roughly a month of that payload into retained data. This is also below 25 GB before assuming compression, but actual storage includes Axiom’s representation and must be measured. This gives a reasonable free-tier design target, not an unconditional promise for any activity level or compression ratio.
Failure/retry sensitivity: in the all-users-active case, sending every batch three times with 10% framing overhead yields 13.84 GB/month at 2% API/bot-evaluation failures, 19.70 GB at 5%, and 29.46 GB at 10%, taking no benefit from repeated-failure coalescing. The heavy-activity case with 2% failures and that same three-transmission assumption reaches 32.90 GB/month. These deliberately combined stresses remain well below ingestion capacity, but their retained-payload proxy exceeds 25 GB in some cases. Logging continues; prolonged incidents may require paid storage. Physical stored bytes and query compute must be measured separately.
Downloader and runner logging must remain tied to summaries and meaningful transitions as data volume and bot counts grow. Per-candle, per-tick, per-contract-success and per-bot-successful-poll logs would invalidate this model. Stress-test event fanout for a 1,000-user workload before production activation; adjust unnecessary event duplication at design time while preserving diagnostic coverage.
Query compute and retained storage
Monitor queries and dashboards consume the independent 10 GB-hours free allowance. Three monitors every five minutes create 25,920 evaluations over 30 days. A planning allocation of 6 GB-hours for those evaluations corresponds to an average 0.83 GB-seconds per evaluation, with remaining allowance for investigation and other account queries. This is a sizing target, not a limit on query execution or a stop rule. Measure actual cost before claiming the free-tier target is met.
Use short query windows, explicit fields and manual dashboard refresh by default. Investigations remain available as needed. Measure retained bytes, actual event size, retry amplification, aggregation cardinality, monitor compute and account-wide usage during activation. Evaluate a continuous 30-day traffic window rather than assuming calendar-month ingestion alone bounds storage.
Growth visibility and decisions
Keep cumulative submitted bytes, accepted/rejected events, retry bytes and per-family event rates as transport observations. Usage-accounting failure is diagnostic and must not disable otherwise functioning log delivery. There is no daily/monthly byte gate, spend gate, budget ledger, reserved byte class or storage/query-triggered export cutoff.
Review actual Axiom account usage and forecasts daily during activation, then weekly once stable. Advisory warning levels at 50%, 70% and 85% of each published allowance prompt investigation; they never stop logging, silence important events or change sampling automatically. Include the projected month-end ingest, steady 30-day retained storage, query compute, and estimated volume at 1,000 daily active users. Account UI/hourly audit observations are the reconciliation source; do not invent an exact storage API.
Keep the three free monitor slots focused on operational failures and exporter health. Usage review/forecasts are part of the existing operational review rather than a fourth Axiom monitor. Check other account datasets/integrations because they share free allowances.
If forecasts approach free capacity before 1,000 users, review duplication, unexpectedly verbose event families, retry loops, broad queries and unrelated integration drains. Preserve important failure and lifecycle information. Paid capacity is an acceptable response when the measured workload requires it; crossing a forecast threshold must not cause an application-imposed logging shutdown. Actual account-plan changes are not part of this spec edit.
Transport reliability settings
The following settings protect process memory, message handling and host disk during outages. They are configurable transport parameters, not log-spending ceilings or fixed allowances on legitimate daily/monthly traffic.
| Setting | Initial proposed value and purpose |
|---|---|
| Structured event | Target under 1 KiB; maximum 2 KiB per record for compact schema/diagnostics. |
| Export batch | At most 50 events or 64 KiB, flush after five seconds; more batches continue as needed. |
| Producer queue | 128 events or 256 KiB per process, nonblocking. |
| Offline spool | 64 MiB logical payload and up to 24-hour age; size from measured traffic/outage tolerance before activation. |
| Gateway state | Initial 256 MiB physical envelope including DB/WAL/diagnostics, adjustable with host capacity. |
| Gateway process | Initial 256 MiB memory allocation and bounded concurrency; benchmark expected traffic. |
| Transport heartbeat | Fresh record every 15 minutes; queued stale heartbeats expire. |
| Routine state summary | One aggregate/service every 15 minutes plus meaningful state changes. |
At the heavy planning workload, 64 MiB holds only about 4.8 hours of unamplified 2 KiB events; the age setting does not promise a full day of retention. Capacity and oldest queue age must be visible, and spool storage can be increased when required. Normal export continues regardless of accumulated Axiom usage.
Keep the spool and aggregation metadata in a dedicated directory, separate from Parquet and Dhan state. Bound SQLite files/WAL growth, checkpoint after small commits, rotate local diagnostics and protect the host’s existing data disk watermark. Queue overflow evicts expired/lower-priority duplicate detail first and reports persistent drop counts. Full disk or spool corruption is a transport fault, not a budget event; preserve native fallback and production operation. No budget-reservation state is needed.
11 Failure handling and acceptance
A gateway 202 means accepted into its durable queue, not accepted by Axiom. Check Axiom’s ingest response counts. Stable event_id is retained through retries; ambiguous network completion can duplicate a record. Deduplicate transport by event_id and causal failures by their authoritative boundary/failure_id. Deduplicate rollup_id before summing its nonoverlapping deltas.
Gateway retry rules: network/5xx/429 use bounded backoff with jitter and Retry-After when supplied; maximum 3 attempts total and 24-hour age. Record retry attempts and observed bytes. Bad schema/data is not retried blindly. Authentication/permission errors stop export until configuration is corrected. Axiom's documented partial-failure response does not provide a reliable event-ID/index mapping. Count and quarantine an ambiguous partial batch; do not resend it wholesale. Retry rejected records only when conclusively identified. Transport summaries retain safe codes only.
Preserve original _time for queued events. A failure arriving more than 10 minutes late is retained for historical investigation and may fall outside the critical alert window. A fresh, deduplicated transport summary reports newly delivered critical backlog counts and original time/age so the critical monitor can raise a delayed-delivery notice. This notice is not a claim that the failures occurred now. A collector, Axiom or notifier outage can still prevent notification; accumulated usage never disables this application-side notice.
The spool prioritizes critical events, evicts expired or lower-priority repeated detail first, and increments persistent drop counters when its physical capacity is exhausted. Increase spool capacity when measured outage tolerance requires it. Missing usage counters do not stop export. There is no application-imposed log-volume or spending cutoff; the next successful transport summary reports any physical delivery loss.
Logging failures do not change product HTTP responses, business results, transactions, provider budgets, or scheduler cadence. Native bounded logs remain the fallback. Gateway acknowledgement and shutdown flushing are best-effort; this system is not lossless.
12 Three monitors and useful investigations
Configure only after implementation and account validation. Start at a 5-minute evaluation cadence, then adjust using measured query compute.
| Monitor | Proposed rule | Interpretation |
|---|---|---|
| Critical operation failure | In the last 10 minutes, an authoritative event has alert_class=critical and a positive newly observed failure count or newly delivered critical-backlog count. | Platform/provider/persistence/corruption failure or explicitly delayed delivery; excludes user validation, cancellation and unverified browser assertions. |
| Work stalled or data degraded | In the last 20 minutes, a current summary/transition has alert_class=degraded, positive rejection/unresolved work, or a violated known progress/freshness threshold. | Includes queue overload, DB discovery failure despite a fresh heartbeat, and known traded-bucket recovery delay. Unknown session completeness is not guessed. |
| Exporter silent | Fresh observability.heartbeat count is below 1 over the last 30 minutes. | Collector/export path may be unavailable or rejected by Axiom. It does not prove every worker is healthy. |
Threshold monitors suit aggregate outcomes; avoid a match alert for every raw error. Three monitors every 5 minutes generate 25,920 evaluations in 30 days, before dashboard use. Their actual compute must be measured. See threshold monitors and match monitor limits.
The gateway keeps an explicit expected-service roster with heartbeat age, owner schedule and declared completion. Missing groups cannot disappear from health summaries and look healthy. Only configured persistent containers/host owners have the 15-minute summary obligation; absence beyond 30 minutes is degraded. Idle Vercel instances and browser clients are excluded. A finite import declared completed is not expected to continue. Scheduled waiting can be healthy without new data. Market freshness uses known session/listing/trade evidence and configured candle delay; with unknown calendar evidence it stays unverified.
Axiom alerts cannot guarantee notification while Axiom, its query allowance, or email delivery is unavailable. Native host/platform diagnostics remain necessary. Test an explicit zero-count heartbeat rule; a grouped query that simply loses a service row is insufficient.
Useful saved investigations:
- Builder failures by stage/provider/reason/release, with repair recovered versus exhausted.
- Engine completed but result save failed, linked by run and operation.
- Requested dataset unavailable versus dependency unverified, with venue and exact contract.
- Published data versus DB checkpoint pending, and oldest unresolved repair.
- Running paper bots versus failed discovery, quote/engine/persistence failures, queue rejection and pending reconciliation age.
- Changes after a release or engine-hash mismatch.
- Export accepted/rejected/suppressed/dropped counts and queue age.
Use interval denominators with their completeness flag and identify whether counts represent collector observations or a complete worker workflow interval. Pre-collector loss makes Next.js/browser totals best-effort. Do not calculate exact product success rates from those incomplete counts or from all failures divided by a 1% success sample.
13 Delivery stages and verification
These stages describe a future implementation, not authorization to begin it.
- Foundation: shared schema/error taxonomy; safe event constructors; gateway, bounded delivery spool, service authentication, usage observations, native fallback and local fixture validation.
- Core failure coverage: builder/provider/check outcomes, backtest lifecycle/save, Data API failures, download publication/bookkeeping, option recovery and paper-runner dependency/engine/queue failures.
- Browser and quality context: restricted browser relay, explicit page/chart failures, tested cross-origin request IDs, auth/credentials, catalog/archive outcomes, expected-owner summaries and container lifecycle metadata.
- Controlled activation: verify account entitlement/other usage, host resources and token scope; enable a small production subset, measure payload/query/storage use for 24 hours, verify the 1,000-user forecast, then widen event families without adding usage cutoffs. Do not launch duplicate workers or new data audits.
- Acceptance: observe one real closed-bar workflow and one backtest/save workflow, verify correlation and compact outcome logs, test the three monitors, and report delivery/drop limitations separately from data correctness.
The staged rollout should use current main in isolated implementation worktrees after approval; the stale original local working trees remain untouched.
Required meaningful checks
| Scenario | Acceptance requirement |
|---|---|
| Builder HTTP 200 error stream and MCP isError | Terminal failure visible with operation/stage even when status remains 200. |
| Provider 429/5xx, context rejection and unknown billed usage | Distinct reasons; recovered versus exhausted; missing usage stays nullable. |
| Engine success followed by DB save failure; secondary usage write failure | Primary and secondary outcomes preserved without altering current product behavior. |
| Real 404, timeout/unverified, malformed metadata, expected empty/session closure | Correct reason and severity; no fabricated completeness or outage. |
| Partial admin/historical backfill | Logs requested/checked/published bounds and partial termination despite current completed/exit status. |
| Parquet publish then DB failure; corrupt file; timestamp collision | Publication, bookkeeping, corruption and normalization separated. |
| Stream connected but pending traded repairs/no authoritative publish | Progress/coverage is distinct from socket liveness. |
| Fresh runner heartbeat with discovery DB failure; queue reject; worker timeout | Dependency/queue/engine failures remain visible and bounded. |
| Thousands of identical contract/bot failures and a 1000-user workload | Aggregation preserves distinct reasons/recovery and occurrence counts; measure fanout, event bytes, retained storage and query compute. No usage level causes export shutdown. |
| Export timeout, partial acceptance, 429/5xx, token rejection | Bounded retries measured; accepted data not repeatedly resent; duplicate IDs handled and transport recovers. |
| Restart, deployment, lost usage counters, full disk and WAL growth | Usage observations resume without disabling delivery; spool/WAL/resource faults have bounded fallback and product continues. |
| Browser abuse and seeded secrets/code/financial payloads | Allowlist and source abuse protection hold without a global daily event cutoff; no arbitrary content in Axiom, spool or alert. |
| Gateway and owner silence, finite-worker completion, scheduled waiting | Explicit expected roster and zero-heartbeat rule; no false healthy empty groups. |
| CORS/range/FileResponse/streaming lifecycle | Response IDs and failure stages verified without breaking data delivery. |
Before claiming completion, verify local focused tests, account receipt/readback, measured production volume and the 1,000-user forecast, Axiom acceptance counts, and a real workflow. No tests or production activation have been performed during this design task.
14 Separate correctness followups
The logging audit revealed issues that observability alone does not repair:
| Concern | Separate change to consider after review |
|---|---|
| Partial backfill written with overwrite then marked complete | Correct outcome/storage semantics and require genuine coverage before completion. |
| Generic health can hide DB failure or incomplete progress | Define readiness and coverage semantics for existing endpoints/UI. |
| Provider exception or primary failure replaced by finalization error | Preserve causes without masking the primary outcome. |
| Legacy browser cache identity lacks venue | Resolve cache identity consistently with data contracts. |
| Reconciliation admission failure currently ignored | Decide execution/backpressure behavior after measuring rejection. |
| Fixed market-session rules are not full calendar evidence | Design calendar/completeness verification separately if stronger guarantees are required. |
Instrumentation may record these paths, but must not silently change downloads, scheduling, persistence, cache behavior, trading decisions, auth responses, or existing API status contracts.
15 Review outcome
Approval of this spec selects a logs-only architecture, the event families/privacy policy, one shared exporter and a measured free-tier capacity target through at least 1,000 product users. It does not introduce a log-budget ceiling or automatic export cutoff. Production setup and implementation remain the next stage.
Reassess current repo heads and actual account/host conditions when that stage begins. Paid logging capacity is an acceptable growth path when measurements justify it. Additional servers, broad drains or stronger data-audit scope are separate design changes. This revision changes the spec only; account upgrades and implementation have not been performed.