Build your online resume. Claim your username
Alpaca logo

Incident Operations Commander Remote - Global at Alpaca

View Alpaca jobs Verified
Remote 🌍 Work from Anywhere Full time Mid Posted  Apply before Oct 23, 2026

Job Description

About Alpaca

Alpaca is a US-headquartered, global leader in agent-first brokerage infrastructure, offering services for stocks, ETFs, options, crypto, fixed income, 24/5 trading, and more. As a licensed financial services company, Alpaca serves hundreds of financial institutions across 40 countries, including broker-dealers, investment advisors, wealth managers, hedge funds, and crypto exchanges, managing over 10 million brokerage accounts. Their global team comprises experienced engineers, traders, and brokerage professionals dedicated to opening financial services to everyone worldwide. Alpaca is committed to open-source contributions, fostering a vibrant community, and continuously enhancing its award-winning, developer-friendly API and robust infrastructure. The company is backed by $400 million in funding from top-tier global investors.

Alpaca's team of over 400 members is globally distributed, thriving from various locations across the USA, Canada, Japan, Hungary, Nigeria, Brazil, the UK, and beyond. They seek passionate individuals aligned with their core values - Stay Curious, Have Empathy, and Be Accountable - who are ready to make a significant impact.

Role

Serve as the on-duty commander for Alpaca's most critical incidents. This role involves directing cross-functional responses to quickly restore service, ensuring the right people are engaged and informed, and guaranteeing that every incident leads to actionable organizational improvements. Your primary responsibility is not to fix the outage but to ensure the response is reliable: assigning correct severity, involving the appropriate engineers, facilitating swift mitigation, keeping leaders informed promptly, and ensuring follow-up work is completed.

Things You Get To Do

  • Command Incidents End to End: Take charge from declaration to mitigation, keeping responders focused on minimizing customer and partner impact. Manage the bridge, protect responders from distractions, and identify when progress is stalling.
  • Classify and Hold the Line on Severity: Set and re-evaluate incident severity. While Risk advises on financial and regulatory materiality, the final decision on severity rests with you.
  • Engage the Right People, Fast: Quickly identify and page the owning team based on service, symptom, and blast radius. Expand the responder group if the initial team is insufficient. Escalate unanswered pages and bring in leaders for critical business decisions (e.g., feature flags, traffic shedding, failover, freeze-or-ship).
  • Hold the Bridge and Protect Responders: Shield engineering and technical support from external inquiries, acting as the single point of contact for stakeholders, partners, and executives. Serve as the sole source of truth for the partner communications team regarding impact, severity, and timing, deciding when status page updates or partner contacts are necessary.
  • Run Follow-the-Sun Handoffs: Provide warm, high-fidelity handoffs across regions, detailing current impact, severity, mitigation paths, next actions, involved personnel, outstanding decisions, and critical information that must not be overlooked. Ensure the incoming commander confirms ownership before you disengage.
  • Close the Loop, On the Clock: Maintain a real time incident timeline for regulatory compliance. After mitigation, ensure a blameless retrospective is scheduled with a named owner and a timebox. Record the cause to determine mandatory follow-up actions. Every action item requires a real ticket, a named accountable person, a priority, and a category, all delivered within agreed service levels. Escalate to SRE if a postmortem yields only low-priority items, signaling incomplete analysis.
  • Automate Coordination Away: Identify and advocate for automating routine coordination tasks. Contribute to decision trees, provide feedback on their accuracy, and articulate the rules behind your instincts to facilitate the transition to AI-driven incident management.

Who You Are (Must-Haves)

  • 4+ years commanding or co-commanding high severity incidents in a production engineering, SRE, or technical operations environment.
  • Ability to direct technical responders under pressure without being the person writing the fix.
  • Proven ability to make and defend crisp severity and escalation decisions, and to take charge instinctively without waiting to be asked. This includes confidently waking senior people at 03:00, interrupting executives, and directing experienced engineers with an audience observing.
  • Proficiency in reading dashboards and independently assessing whether impact has stopped.
  • Clear communication skills with engineers, executives, and partner facing stakeholders, understanding the distinction between briefing communications and speaking for the company.
  • Comfortable holding other teams accountable in the moment, across different reporting lines, without creating friction.
  • Thrives in a follow the sun model with clean cross region handoffs.
  • Understanding of FinTech concepts and the high trust stakes of API driven financial platforms.
  • Experience using AI tools and agentic automation to reduce manual toil and speed up incident response.
  • Willingness to work a regional coverage window as part of a global 24x7 Incident Commander roster.

Who You Might Be (Nice-to-Haves)

  • Formal incident command training (e.g., ITIL, Major Incident Management, crisis management).
  • Experience with modern incident management and on call platforms.
  • Previous experience writing severity rubrics, decision trees, escalation matrices, runbooks, or incident playbooks.
  • Commanded incidents in game days, tabletop exercises, or simulations, not only in production.
  • Partnered with problem management or reliability program functions to roadmap incident follow ups.
  • Online securities trading or capital markets experience, or experience in another regulated, market hours sensitive domain.

How We Take Care of You

  • Competitive Salary & Stock Options
  • Health Benefits
  • New Hire Home Office Setup: One time USD $500 stipend
  • Monthly Stipend: USD $150 per month via a Brex Card

Alpaca is an equal opportunity workplace dedicated to pursuing and hiring a diverse workforce.

Ready to Apply?

Take the next step in your career journey.

Apply Now

You will be redirected to the company's application page

Link verified about 1 hour ago

💜 Please mention that you found the job on True Work From Home, this helps us grow. Thanks!