Evidence ruleThe report proves heap exhaustion at this scale. It does not distinguish a leak from legitimate retained workflow state without per-batch heap telemetry.

Two limits answer different questions.

Live overlap0 → up to 16

maxConcurrentAgents defaults to auto: min(16, max(1, cores - 2)).

Run lifetime1000 total

maxTotalAgents stops excessive calls. It is not a heap budget.

Turn one long fan-out into restartable work.

  1. Partition the input. Make every batch independently restartable.
  2. Start at 2-4 concurrent children. Raise one dimension only after measurement.
  3. Set a batch-sized total ceiling. Do not inherit 1000 when the batch expects 64 calls.
  4. Persist before the next batch. Keep completed output outside the workflow lifetime.
  5. Record memory and outcomes. Capture peak heap, child count, duration, failures, and retries.

Use a bounded delegation group.

- id: delegation
  name: cordis:group
  group: true
  isolate:
    workflows: true
  config:
    - id: workflow-worker-thread
      name: '@deepseek-ai/dsh-workflow-worker-thread'
      config:
        provider: spawn
        maxConcurrentAgents: 4
        maxTotalAgents: 64
    - id: tool-workflow
      name: '@deepseek-ai/dsh-tool-workflow'

The values are conservative starting points, not universal sizing advice. Increasing --max-old-space-size only moves the failure boundary if the workload still retains too much state.

Stop on the memory trend.

Heap falls after a batch

Continue cautiously. Compare the next batch at the same bounds.

Heap rises after settled children

Stop growth. Preserve a heap profile and the settled-child count.

OOM before total limit

Reduce both ceilings and shorten the process lifetime.

Primary evidence.

Keep the runtime boundary visible.

Use the live status board for source-backed reports, first evidence, and safer next actions.

Open field status