Worfilo
A hosted editor for AI agent workflows: describe one or wire it on a canvas, watch every node run live, replay any run, and ship it as an API.
Solo designer and engineer · Sep 2026 - Present · Live, early access
Worfilo is a visual builder for AI agent workflows, and its pitch is three verbs: describe a workflow, watch it run, ship it as an API. Nodes sit on a canvas, get connected and configured, and run with live per-node status. A workflow can also be drafted from a plain-language description, which Claude turns into a validated graph that the builder previews before creating it.
It is built for developers and technical builders who are comfortable reading JSON and writing templates such as {{ nodes.classify.output.structured.urgency }}, and who want agents wired into working automations faster than hand-written orchestration code allows.
I was the only person on it: product, UX, the React Flow editor, the FastAPI engine, the AWS infrastructure, the docs and the marketing site. That is why so many decisions below favour fewer moving parts.
- 14
- node types, the whole kit
- 5
- built-in connectors, plus any HTTP API or MCP server
- 175
- backend tests, plus a Playwright suite
- 1
- JSON graph behind every workflow and every run
The problem
General automation tools treat AI as one more integration. For agent work that causes three recurring pains.
Agents are bolted on. Tool use, structured output and prompt iteration feel like workarounds.
Failed runs are opaque. When an agent gives a wrong answer the question is always which step went wrong and with what input. Logs answer that slowly, especially after the workflow has since been edited.
Catalogues are bloated. Hundreds of integrations bury the few that matter and widen the security surface.
The approach
Worfilo makes four commitments and lets every design decision follow from them.
Agent-first
The agent is the core node. A tool edge attaches any capable node to an agent as a callable tool, structured output lets the agent return JSON that drives a branch, and a single node can be re-run against the last run's data to iterate on a prompt.
Debuggable
Every run stores the exact graph it executed. A finished run replays on a read-only canvas, and each node's recorded input and output stay inspectable after the workflow changes.
Small and understandable
Fourteen node types and five connectors (GitHub, Slack, Notion, Linear, Tavily). Custom APIs and MCP servers cover everything else.
Private by default
Keys and tokens are encrypted at rest, write-only, never placed in a workflow, and scrubbed from run records. A Mask PII node keeps personal data away from the model.
The central idea behind the architecture: the canvas is only an editor for a JSON graph. The backend never knows about pixels, it receives a graph, validates it and executes it. The frontend never executes anything, it renders the graph and subscribes to run events.
How a run works
- 1
Describe or draw
Write a sentence and review the generated graph, or wire nodes by hand. Either way the result is the same JSON graph.
- 2
Validate
The same checks run on both paths: configs match their schemas, templates read only upstream nodes, there are no cycles, and every node is reachable from the trigger.
- 3
Run
The engine executes the graph and streams each node's status, output and tokens to the canvas over Server-Sent Events.
- 4
Replay
The run keeps a snapshot of the graph it executed, so it reopens later on a read-only canvas with every input and output intact.
- 5
Publish and call
Publishing freezes a numbered version. Triggers and the workflow API run that version while the draft keeps changing.
Key features
Visual editor and live runs
A React Flow canvas with a searchable node palette, schema-driven configuration forms, draft autosave and validation as you type. Runs stream per-node status over Server-Sent Events, with inputs, outputs and tokens inspectable as they happen.
Run replay and node testing
Because each run snapshots its graph, any past run can be reopened on a read-only canvas and inspected node by node. The branch that was not taken stays marked skipped.
Test this node runs a single node against the last run's upstream outputs, with the configuration on screen whether it is saved or not. That makes prompt iteration cheap.
AI generation
Describe a workflow and Claude drafts the whole graph against the live list of node types, using your connectors, APIs and MCP servers by their real ids. The draft goes through the same validation as a hand-built workflow, with one repair round, and is previewed on a canvas. You can ask for changes in plain language before it is created.
Sonnet 5.5 gives better graphs and Haiku 4.5 gives faster, cheaper drafts. Generation runs on Worfilo's own key and is rate limited per account.
Integrations
Connectors: GitHub, Slack, Notion, Linear and Tavily, by token or OAuth. Tokens are checked with the app before they are saved.
Custom APIs: any HTTP API with no auth, basic, bearer, API key, custom headers, digest, OAuth 2.0 (client credentials and authorization code with PKCE) or AWS Signature v4. Saved endpoints turn {placeholders} into node inputs.
MCP servers: remote Model Context Protocol servers, added from a marketplace backed by the Smithery registry or by URL. One tool becomes a node, or a whole server attaches to an agent. Servers that run as a local command are not supported, because Worfilo runs on a server, not on the user's machine.
Execution engine
Branching with If and Switch, merging, parallel branches, retries with backoff, per-node and per-run timeouts, a per-run token cap and three error modes (stop, continue, fire an error port). Templates may reference only upstream nodes, and validation rejects anything else before a run starts.
Triggers, publishing and the workflow API
Publishing freezes the draft into an immutable numbered version, so editing never changes what production runs. Schedule triggers live on their own page with cron expressions, timezones and a preview of the next three runs, and always run the published version.
A published workflow can also be called from code. API keys can be scoped to chosen workflows, are shown once and are stored only as a hash. Callers can wait for the result, poll a run or stream it over SSE, and each workflow has a playground with ready-to-copy cURL and Python.
Built for coding agents
One command, npx @worfilo/mcp install, adds the Worfilo MCP server to Claude Code, Cursor, VS Code, Windsurf and Antigravity. The agent can discover node types and connections, plan a graph, create and test a draft, publish it, and create a scoped API key. It signs in over OAuth 2.1 with PKCE and asks for narrow permissions such as read, write, publish and run, which can be revoked at any time. The product also ships llms.txt and an AI catalog for discovery.
PII Shield
A standalone service, exposed as the Mask PII node, that replaces personal data with tokens such as <PERSON_1> before text reaches a model and restores it afterwards. Three detectors run together: recognizers with checksum validation for structured identifiers, a multilingual GLiNER model for free text, and a caller-supplied always-mask list. Profiles (General, Healthcare, Finance, Minimal) and per-country identifiers, including Sri Lankan national ID numbers, tune what is masked.
Original values are Fernet-encrypted in a short-lived store and never written to disk, and a run records only how many values of each type were masked. If the shield is down the node either passes text through or fires its error port, whichever the builder chose.
Architecture
Browser -> Cloudflare proxy -> app host (EC2, Docker Compose) caddy -> web (Next.js) | api (FastAPI) | worker (arq) | pii-shield | valkey (queue, vault) RDS PostgreSQL 17, private subnets, point-in-time restore
The executor has no special case for any node type. It follows one rule: an outgoing edge is active only if its source port appears in the node's result, and dead otherwise. A node whose incoming edges are all dead is skipped, and its own edges die too.
An If node fires true or false, so one branch dies. A node that fails in continue mode fires nothing, so everything downstream is skipped. A Merge waits until every incoming edge is settled. Branching, skip cascades and error ports all fall out of that rule, which is also why the canvas can explain every skipped node.
Every step is an event with a sequence number within the run. If the connection drops, the editor reconnects with the last sequence number it saw and receives exactly the events it missed, none replayed twice and none lost.
| Frontend | Next.js App Router, TypeScript, React Flow, Zustand, shadcn/ui, Tailwind |
| API | FastAPI, Pydantic v2, SQLAlchemy 2, Alembic |
| Engine | Async executor with a node registry, run context and event bus |
| Templating | Jinja2 sandboxed environment, strict undefined |
| Realtime | Server-Sent Events with resumable sequence numbers |
| Queue | arq on Redis in production, in-process asyncio locally |
| Auth | Supabase Auth, tokens verified against the project's public signing keys |
| LLM | Anthropic, behind a provider layer built to accept others |
| Secrets | Fernet encryption with key rotation |
Decisions and trade-offs
A JSON graph, not a code-first engine. It costs flexibility, since graphs are acyclic and there is no loop node yet. It buys validation, replay, generation and a stable API from one data model.
Five connectors, not five hundred. Everything else is an HTTP API or an MCP server. The catalogue stays readable and the security surface stays small.
Snapshot the graph on every run. It stores more data per run, but a failed run can always be explained against what actually executed.
Generation reuses validation. AI drafts go through the same checks as hand-built graphs, so there is no second and weaker path into the engine.
Triggers run published versions only. A schedule keeps running what was last published while the draft changes, so editing is never a production event.
| Node timeout | 120 seconds |
| Attempts per node | 1, retries back off from 2 seconds |
| Parallel branches | 4 at a time |
| Run timeout | 900 seconds |
| Token cap per run | 200,000 |
| Fastest trigger | every 5 minutes |
Security and operations
SSRF protection. HTTP and API nodes refuse private, loopback and link-local addresses and non-HTTP schemes. A custom API's path cannot leave its base host, and redirects are not followed.
Secrets. Write-only credentials, redacted from run records (including secrets an API echoes back), OAuth tokens refreshed on the server, key rotation script included.
Edge. Port 443 accepts only Cloudflare's addresses and Caddy rejects requests missing a secret origin header.
No SSH. Shell and database access go through SSM Session Manager.
Infrastructure as code. Terraform for network, EC2, RDS, ECR, DNS, alarms and audit, deployed from GitHub Actions over OIDC.
CI. Lint, type checks, tests, migration checks, dependency audits, gitleaks history scan, Trivy and SonarQube quality gates, with compliance evidence uploaded to CISO Assistant.
Design
A quiet, legible interface for technical users. Geist for text and Geist Mono for code, ids and data. Motion carries meaning only: it signals an agent thinking, a node executing or a production-grade action, and is never decoration.
Status is never conveyed by motion or colour alone, every state also has an icon and a text label. The marketing page demonstrates the product with a scroll-driven demo run and a replay, rather than screenshots.
Challenges
Keeping the graph contract stable. The frontend saves it, the backend validates and executes it, and every run snapshots it, so the data model was designed first and everything else hangs off it.
Safe templating between nodes. Sandboxed Jinja2 with strict undefined values, plus upstream-only reference validation, keeps expressions powerful without letting a run reach Python internals. A misspelled field fails loudly instead of rendering empty.
Trust in agent runs. Replay, per-node testing and recorded inputs and outputs exist so that a failed run can always be explained.
Scheduling that fires once. Each trigger is claimed and moved to its next slot before its run starts, so it fires once even with several workers, and a scheduler that was down for an hour fires once on return, not sixty times.
Shipping production hardening solo. Edge lockdown, secret handling, PII masking and compliance evidence were built alongside the product rather than after it.
Outcomes
Verified live against the real Anthropic API: a Haiku classifier returning structured {category, urgency} that drives an if branch, and an agent that chose its own arguments and called the live GitHub API through a tool edge.
Worfilo is live in early access at worfilo.com, with documentation for every node, connector and the workflow API, seven worked use cases and a blog covering the engine design. The project has 175 backend tests and a Playwright end-to-end suite.
What is next
Webhooks. The webhook trigger node exists, and the public endpoint that fires it is next.
Paid plans. The free plan limits workflows, concurrent runs and runs per day. Larger plans are coming.
More providers. Additional LLM providers behind the existing provider layer. Agent nodes use Anthropic today.
Loops. A dedicated loop node, since graphs are acyclic by design today.
Scale. Horizontal scaling across workers.
Stack
- Next.js
- TypeScript
- React Flow
- Tailwind CSS
- FastAPI
- Python
- PostgreSQL
- Redis
- Supabase
- Claude API
- MCP
- AI Workflows
- Docker
- Terraform
- AWS
- GitHub Actions
- Playwright