Durable execution for background work
Write it once. Never babysit it again.
You wrote the job in an afternoon. Then you wrote the retry table, the status column, the sweeper that finds the stuck ones, the runbook for when it gets paged. Hopskip is the second half of that work, done once, by the engine: your function survives the machine it runs on, and every step happens exactly once.
Survives kill -9, OOM kills, deploys, and lost regions.
Java and .NET workflows are not supported. Language support.
I
The infrastructure you'd otherwise write.
The table with the status column. The poller that re-enqueues the stuck rows. The idempotency keys threaded through every call by hand. Every team builds this eventually, badly, and then maintains it forever. Hopskip is that layer, built once, correctly: each step's result commits to a log that outlives any machine. Kill the worker mid-step and another one resumes from the last commit. No half-written rows, no orphans, no sweeper.
- no retry loop with a counter column
- no status table nobody trusts
- no cron sweeper for stuck rows
V
The deploy isn't a migration project.
Add a step and ship it. The deploy gate checks every in-flight run against the new version, carries each one to its matching step, and refuses the ones that can't move safely. No drains, no dual-deploy windows, no version-branching checks left to rot in your business code.
- the new version is checked before it goes live
- running workflows move to matching steps
- unsafe changes are blocked, not half-applied
III
The language your team already writes.
Write workflows in TypeScript, Python, Go, Rust, or Haskell. The five SDKs share a workflow contract: the engine records activity outcomes and replays them without repeating external work. Activities can call the services your team already operates.
- sandboxed workflow execution
- activities can call existing services
- a shared replay contract across five SDKs
IV
Debug what ran. Not what you guess ran.
Every await point is a snapshot you can restore. Scrub to the step that failed, open it with your real variables in memory, and fork it to test the fix against the actual inputs. Asking a run a question is a lookup, not a search cluster you operate on the side.
- restore any step, attach a real debugger
- query live state, no separate visibility stack
- test the fix against the exact failing inputs
II
Kill the worker. Keep the work.
Step through it. A worker is killed mid-run; another claims the work, restores the last commit, and finishes. Every step ran exactly once, and nobody wrote a recovery script.
- resumes from the last step, not from zero
- each step's side effects happen once
- same code in dev, staging, and production
Run it yourself, or let us.
Managed in our cloud, licensed for your infrastructure, or free for evaluation and small workloads. The rates below are the ones on the pricing page.
Hopskip Cloud
fully managed · early access
No base fee
pay for compute, storage & network
- metered only while code runs
- published rates for AWS and Google Cloud
- never billed per step or per event
- managed upgrades, recovery, and capacity
Self-hosted
commercial · your infrastructure
from $2,500/mo
dev, staging, and cold standby free
- commercial license for production use
- S3, tiered cold storage, multi-region
- OIDC, SAML, SCIM, and the compliance pack
- first 16 hopskip-server vCPUs included
Community
evaluation and small deployments
$0
forever · community support
- sandbox, event store, dispatch, snapshots, replay
- all five SDKs, the console, SQL visibility
- clusters of up to three nodes
- local disk, bearer-token auth