Senior Site Reliability Engineer at Alpaca
Job Description
About Alpaca
Alpaca is a US-headquartered, global leader in agent-first brokerage infrastructure, offering comprehensive services for stocks, ETFs, options, crypto, fixed income, and 24/5 trading. Our licensed financial services subsidiaries cater to hundreds of financial institutions across 40 countries, including broker-dealers, investment advisors, wealth managers, hedge funds, and crypto exchanges, collectively serving over 10 million brokerage accounts.
Our diverse global team comprises 400+ experienced engineers, traders, and brokerage professionals dedicated to our mission of opening financial services to everyone worldwide. We are deeply committed to open source contributions, continuously enhancing our award winning, developer friendly API and its robust infrastructure. Alpaca is proudly backed by $400 million in funding from top-tier global investors like Portage Ventures, Spark Capital, Tribe Capital, Social Leverage, and Y Combinator.
We are a dynamic, globally distributed team with members spanning the USA, Canada, Japan, Hungary, Nigeria, Brazil, the UK, and beyond. We seek passionate individuals aligned with our core values - Stay Curious, Have Empathy, and Be Accountable - who are ready to make a significant impact on Alpaca's rapid growth.
The Opportunity: Senior Site Reliability Engineer
As a Senior Site Reliability Engineer at Alpaca, you will be instrumental in maintaining and enhancing the reliability, observability, and operability of our critical brokerage platform as we continue to scale. Your work will span our cloud infrastructure, Kubernetes platform, observability stack, messaging layer, and data layer. We are particularly interested in candidates with strong PostgreSQL fundamentals who are eager to develop deeper ownership of our database reliability posture. PostgreSQL is critical to our trading path, and this role involves a meaningful share of time dedicated to its enhancement, alongside broader SRE responsibilities.
What You'll Do
- Operate production systems day-to-day, including oncall duties, incident response, postmortems, and follow ups to ensure issues are resolved comprehensively.
- Own reliability practices, defining and refining SLIs, SLOs, and error budgets, and guiding product teams to operate within these parameters.
- Strengthen our observability capabilities across metrics, logs, traces, and alerting systems.
- Ship infrastructure through code using a GitOps workflow, encompassing both cloud resources and Kubernetes workloads.
- Look after PostgreSQL, focusing on performance tuning, schema and migration review, online migrations on large tables, high availability (HA) and disaster recovery (DR) strategies, and Change Data Capture (CDC) pipelines.
- Mentor engineers on reliability best practices and database fundamentals through code review, design review, and pairing sessions.
What You'll Bring (Must Haves)
- 4+ years of professional experience in Site Reliability Engineering (SRE), DevOps, Platform Engineering, Infrastructure Engineering, or backend engineering roles with significant production operations ownership.
- Hands on experience operating production services on Kubernetes and deploying infrastructure as code within a GitOps workflow.
- Solid working knowledge of PostgreSQL in production environments, including understanding query plans, pg_stat_* views, indexing and schema trade offs, and safe online migration techniques for non trivial tables.
- Strong understanding of cloud networking fundamentals, such as VPCs, routing, L4/L7 load balancing, DNS, and TLS, with comfort in debugging cross-service connectivity issues.
- Comfortable with modern observability stacks and proficient with Linux at an operator level.
- Practiced in incident response, demonstrating calmness under pressure, structured debugging approaches, and conducting postmortems that drive meaningful change.
- At least working proficiency in Go or Python, coupled with strong written and verbal communication skills.
- A genuine interest in databases and a desire to grow your PostgreSQL and DBA expertise.
Bonus Points (Nice to Haves)
- Deeper PostgreSQL experience, including managing large clusters under OLTP load, performing online migrations on big tables, ownership of HA/DR solutions, connection pooling at scale, or implementing Change Data Capture (CDC) pipelines.
- Experience with typed SQL access layers in Go, such as pgx, gorm, or sqlc.
- Production experience with messaging systems at scale, including RabbitMQ, Kafka, or Redpanda.
- Security and compliance experience within a regulated environment (e.g., SOC 2, secrets management, audit logging).
- Familiarity with trading, brokerage, or other regulated fintech domains.
Our Commitment to You
- Competitive Salary & Stock Options
- Comprehensive Health Benefits
- New Hire Home Office Setup Stipend: One-time USD $500
- Monthly Stipend: USD $150 per month via a Brex Card
Alpaca is proud to be an equal opportunity workplace, dedicated to pursuing and hiring a diverse workforce.
Ready to Apply?
Take the next step in your career journey.
Apply NowYou will be redirected to the company's application page
Link verified about 12 hours ago
💜 Please mention that you found the job on True Work From Home, this helps us grow. Thanks!
More Software Development Engineer (SDE) Jobs
Discover similar opportunities that match your skills
Platform Security Engineer
IT Engineer - Remote
Unity Technical Artist for Immersive Events
Engineer to own the automated pipelines
Backend Engineer, Security (Remote)
Senior Infrastructure Security Engineer
Engineering Team Lead
Senior Software Engineer, Core Services
About Alpaca
Alpaca provides a developer first API platform for trading stocks, ETFs, options, and cryptocurrencies. It enables builders to embed investing features into their applications with commission free access and seamless infrastructure.
View Company Profile