Senior DevOps / Infrastructure Engineer (Americas, US time zones)
Remote (Americas, US time zones) · Full time
About Startale
Startale builds onchain finance as a vertically integrated stack: the chains, a JPY stablecoin, payments, and a consumer app. With SBI Group we are building Strium, a fully onchain exchange; we issued JPYSC, Japan's first trust-licensed JPY stablecoin; with Sony Group we operate Soneium; and Startale App, used by 210k monthly users, is the front door to all of it. We raised a ~$63M Series A from Sony, SBI and others. The team is ~70 people across 20 nationalities, and English is the working language of engineering.
Why US time zones
Our engineering team is spread across Asia and Europe. This role anchors infrastructure coverage in US hours: you will be the infrastructure owner awake when most of the team is not. You must be physically based in a US time zone (UTC-8 to UTC-5), whether in the US, Canada, or Latin America. Working US hours from another region does not qualify.
About the Role
We are looking for a Senior DevOps / Infrastructure Engineer to own the infrastructure, deployment, and operational systems behind StartaleApp, Strium, and other products across the Startale ecosystem. This is a senior, hands-on role: you design, build, and operate production infrastructure end to end, from cloud and Kubernetes to CI/CD, databases, observability, and security.
You will work primarily on two demanding products. StartaleApp is a user-facing Web3 wallet and application platform whose backend combines traditional cloud services and databases with blockchain nodes, indexers, and transaction processing. Strium, built with SBI Group, is a next-generation decentralized exchange bringing Japanese equities onchain, with a fully onchain order book and a consensus layer purpose-built for low-latency, high-throughput trading.
Both run under high load, strict availability and security requirements, and real production conditions. We want someone who builds infrastructure rather than maintains it: comfortable making architectural decisions, automating everything that repeats, and diagnosing complex production issues. Experience running blockchain infrastructure is a significant advantage, and a SecOps or cloud security background is a major plus.
What You Will Do
- Own infrastructure architecture and operations for StartaleApp, Strium, and other production systems.
- Design, build, and operate highly available cloud infrastructure on AWS: Kubernetes clusters, networking, storage, scaling, upgrades.
- Manage everything as code: Terraform, GitOps, and CI/CD pipelines with safe rollouts and rollback.
- Operate PostgreSQL in production: availability, replication, backups, performance, recovery.
- Build observability across infrastructure, applications, databases, and blockchain nodes, with alerting that catches incidents early.
- Run blockchain infrastructure (nodes, RPC services, indexers) and monitor it down to protocol level: sync, connectivity, transaction processing.
- Design disaster recovery and failover for critical systems, and harden security: access control, secrets, vulnerability management, monitoring.
- Diagnose production incidents across the full stack, and remove operational bottlenecks with automation instead of manual procedure.
What You Bring
- 5+ years in DevOps, SRE, platform, or infrastructure engineering.
- A track record of building and operating production infrastructure for highly available systems, where availability directly affected users or the business.
- Deep hands-on Kubernetes, AWS, and Linux experience.
- Infrastructure as code with Terraform, GitOps practices, and CI/CD pipeline design.
- Solid networking fundamentals: DNS, load balancing, TLS, service discovery.
- PostgreSQL in production: replication, backups, recovery, performance.
- High-availability design, incident response, root-cause analysis, and capacity planning.
- Hands-on experience operating blockchain infrastructure (nodes, RPC, indexers, validators, or data pipelines) and an understanding of how distributed blockchain networks fail.
Strong Plus
- Security: SecOps or cloud security experience, hardening production environments, AWS security services and IAM design.
- Platform breadth: bare-metal operations, Helm or ArgoCD, messaging systems such as Kafka or NATS, observability stacks such as Prometheus, Grafana, Loki, or OpenTelemetry.
- Blockchain depth: operating chain infrastructure at significant scale; Ethereum, Polkadot, or L2 ecosystems.
- Domain: financial, trading, or other latency-sensitive systems; using AI tools in infrastructure workflows; Go, Rust, or TypeScript.
How We Think About This Role
We care less about specific tools and more about how you engineer:
- Build infrastructure, not just operate it, and own problems from architecture through production.
- Design for high load and partial failure, and automate what repeats.
- Diagnose across the full stack, from Kubernetes and networking to databases and blockchain nodes.
- Treat infrastructure as a product, with users, reliability requirements, and long-term maintenance needs.
Why Join
- Build and operate the core infrastructure behind StartaleApp and Strium, including a DEX designed for high-throughput, low-latency trading.
- One role spanning cloud, Kubernetes, distributed systems, databases, and blockchain infrastructure.
- Significant ownership and direct product impact in a small, highly technical team.
- Competitive compensation and equity.