troubleshooting-flows

Diagnose and resolve Celigo flow failures -- total failures, partial errors, stuck jobs, empty runs, and performance issues. Use when a flow is failing, producing errors, returning no data, or running slowly.

By celigo · 1,046 installs

npx skills add celigo/ai --skill troubleshooting-flows

Source repository · Upstream listing

<! TIER:1 Troubleshooting Flows A flow is broken when it fails to move data correctly. This skill covers systematic diagnosis: identifying the problem type, isolating the failing step, inspecting errors, and resolving them. Troubleshooting concerns: Job status understanding what completed , failed , canceled , and retrying mean for the flow Error analysis grouping errors by pattern to find root causes instead of reading them one by one Request/response inspection seeing exactly what was sent and returned at each step Execution logs record level tracing through every stage of the pipeline Retry and resolution fixing error data and retrying vs bulk resolving Delta/state issues lastExportDateTime drift, stuck deltas, re processing windows Problem Categories Total Failure Job status is failed with 0 successful records. The entire run collapsed before processing any data. Typically numPagesGenerated: 0 (export level failure) or pages generated but numPagesProcessed: 0 (import level failure on the first page). Common causes: connection failure (credentials expired, endpoint down), export query error (invalid SQL, bad saved search ID), missing/deleted resource, permission denied. Partial Failure Job status is completed but numError 0 alongside successful records. Some records failed while others processed normally. Real world data shows wide variance from 2 errors in 8000 successes to 500+ errors in 8000 successes. Common causes: validation errors on the destination (required fields missing, type mismatches), duplicate key violations, record level lookup failures, rate limiting on specific batches, data dependent issues (specific records have bad data). Empty Run Job completes successfully with 0 errors AND 0 records processed. Common causes: wrong resourcePath on the export (extracts from wrong JSON path), delta export with no changes since last run (legitimate), output filter too restrictive (all records filtered out), source query returns no results, webhook export with no inbound events. Stuck or Long Running Job stays in running or retrying status longer than expected. Common causes: large dataset with no pagination limits, destination system slow to respond, script hook with long running logic, on premise agent connectivity issues, rate limiting causing backoff. Intermittent Failures Flow sometimes succeeds and sometimes fails with the same configuration. Common causes: token/session expiry mid run (long running flows), rate limiting (varies with concurrent flows), transient network errors, source system maintenance windows. Error Diagnosis Framework Classification When an error occurs, classify it into one of three categories to determine the right action: Category HTTP status codes Meaning Action Needs investigation 400, 401, 403, 404, 405, 409, 422 Missing info, wrong IDs, permission denied, validation errors Stop and investigate check resource config, connection status, permissions Transient 408, 429, 500, 502, 503, 504 Timeouts, rate limits, server errors Retry once. If it fails again, escalate the external system may be down Configuration error varies Preconditions not met but fixable Follow the error message guidance to fix the config, then retry A 5xx error not in the transient list (e.g., 501) is still likely transient. A 4xx error not in the investigation list warrants manual review. Root Cause: Configuration vs Data Every flow error has one of two root causes: Static configuration a hardcoded value in the step config is wrong (mapping expression, filter rule, hardcoded field, URI, query, SQL statement). Fix: change the resource configuration via celigo <type set Dynamic data the upstream source sent unexpected data (missing required field, wrong type, null where a value is expected, unexpected array/object shape). Fix: add input filtering or validation upstream, or fix the source system To distinguish: check if the error reproduces with different input records. If the same error occurs for every record, it's configuration. If only some records fail, it's data. Which Step Failed? Error location determines which resource and skill to investigate: Error location Resource to check Skill Export / page generator Export config (connection, query, resourcePath) configuring exports Import / page processor Import config (mapping, destination fields, operation) configuring imports Script hook Script code (preSavePage, preMap, postMap, postSubmit) writing scripts Mapping Mapping expression (field paths, lookups, hardcoded values) writing mappings Filter Filter expression (s expression syntax, field references) configuring filters Connection Connection config (auth, URL, credentials) configuring connections Quick Reference Symptom First Command Symptom Run first Then Flow totally failed celigo jobs list flow <flowId limit 1 Check numPagesGenerated if 0, export failed; check connection and query Partial errors celigo flows error summary <flowId celigo flows error analysis <flowId <stepId to find root cause pattern Empty run (0 records) celigo jobs list flow <flowId limit 1 Check export config ( resourcePath , delta state, output filter) Stuck / long running celigo jobs current flow <flowId Check job status; if retrying , inspect rate limiting or connection issues Intermittent failures celigo jobs run stats flow <flowId Compare failing vs passing runs; check token expiry and rate limits Silent logic bug (no errors, wrong output) celigo flows test run <flowId export <genId If test run can't reach it, enable execution logging (§6) and run for real Production incident (real traffic matters) celigo flows enable execution logs <flowId then run Read per record I/O with query execution logs / execution log detail ; the failing stage names the problem Key Diagnostic Commands Related Skills [building flows Quick Reference](../building flows/SKILL.md quick reference) flow structure, topologies, and configuration [configuring exports Quick Reference](../configuring exports/SKILL.md quick reference) export configuration and adaptor types [configuring imports Quick Reference](../configuring imports/SKILL.md quick reference) import configuration and adaptor types [writing scripts Quick Reference](../writing scripts/SKILL.md quick reference) script hook debugging and data shapes <! TIER:2 Diagnostic Workflow 1. Check the job status Start with the most recent job to understand what happened. Key fields: status , numError , numSuccess , numIgnore , numPagesGenerated , numPagesProcessed , startedAt , endedAt . A failed status with numPagesGenerated: 0 means the export itself failed don't look at import errors. completed with numError 0 means partial failure at the record level. 2. Get the error summary See which steps have errors and how many. This returns per step error counts. Focus on the step with the most errors first. 3. Analyze error patterns Group errors by message pattern to find the root cause instead of reading them one by one. If most errors share the same message, that's your root cause. Multiple distinct patterns may indicate multiple issues. 4. Inspect individual errors Once you know the pattern, look at specific records and the HTTP request/response that produced the error. flows error request detail is the most powerful diagnostic it resolves the error's reqAndResKey for you and shows exactly what HTTP request was sent and what the destination responded with. (If you already hold a reqAndResKey from debug requests , use debug request detail instead.) 5. Use test runs for safe iteration (try this first) Test runs process a single page without affecting production data or delta state. Fast, safe, no arming, no side effects answers most logic questions. test run returns {metadata, flowJob, childJobs} metadata lists stage names per bubble. Follow with test run step results to get stages[] = [{name, input, output, errors}] per bubble. Errors include retryData inline. Test runs don't advance the delta timestamp you can repeat them safely against the same data. Limitations: imports don't actually submit, mock data is shared across lookups/imports, and some adaptors don't run in test mode. When these bite, escalate to §6. 6. End to end debugging with execution logs (silent/logic bugs, production incidents) When test run can't answer it imports must actually submit, destination behavior matters, or it's a production incident arm execution logging, run the flow for real, then read the per record I/O the run captured. Disable debug when you're done. Each per record log entry names the stage that produced it (matching the test run step results stage shape: { name, input, output, errors } ), so the failing stage tells you where the record broke. For raw HTTP at a bubble, use debug requests / debug request detail (§7). Errors carry retryData inline see §8 to fix and retry. 7. Low level debug primitives (surgical control) These are the debug primitives the §6 workflow builds on use them directly when you want manual control over arming, clearing, or probing: Stage names (for execution log detail stage ): Built in: apiCall , transformation , mapping , inputFilter , outputFilter , responseMapping , responseTransformation , routing Script hooks: the function name wired on the bubble (e.g., preMapHook , postSubmitHook , branchingHook , preSavePageHook , postResponseMapHook ) Test run uses a different stage vocabulary than live /logs/data/query : request / response / parse (three stages) instead of live's merged apiCall ; transformTwoDotZero instead of transformation ; responseMap instead of responseMapping ; router instead of routing . Everything else matches. The live execution log commands (§6) use the live vocabulary. 8. Fix and retry (or resolve) Fix the configuration , then retry: Fix the data when specific records have bad values: Resolve without retry when errors are expected or not worth reprocessing: 9. Verify the fix Run the flow again and confirm clean execution. Monitoring Lenses and Execution Metrics Before diagnosing, pick the lens that matches the question all execution state comes through one of two: Running jobs executing right now. The live view: in progress jobs with real time progress (records processed so far, errors accumulating, pages generated, whether the export phase is done). Use it for "what's happening this moment" a long running job you're watching, or a flow you just kicked off. Reach it with celigo jobs current flow <flowId . Completed historical aggregate per flow. The rear view: run counts, average runtime, success / error / ignore totals, open errors, and when it last executed or errored. Use it for "how have things been going" slow flows, error prone flows, "did the order sync run today." Reach it with celigo jobs list flow <flowId and celigo jobs run stats flow <flowId . Quick test: now running; recently / over time completed. Reading the Execution Metrics Keep the record level counts straight so you don't misread a run: Success, error, and ignore are three distinct outcomes per record. numSuccess processed cleanly; numError failed; numIgnore is not an error it's a record the flow intentionally skipped (filtered out, or a no op upsert). Don't fold ignores into error counts: a run with a high numIgnore and numError: 0 is healthy, not broken. Open vs resolved errors. Open errors are the failures still unresolved a