Working within Swan’s Core Infrastructure team, you will help ensure the reliability, scalability, security, and performance of the platforms that support our financial services. You will take ownership of well-scoped services and operational incidents, improve observability, automate repetitive tasks, and collaborate closely with development, product, and security teams.
This is an independent engineering role for someone who has developed solid operational foundations and is ready to take greater ownership of production systems. You will contribute to incident response, infrastructure improvements, service design reviews, and the continuous improvement of our reliability practices.
Main responsibilities
On a daily basis, you will:
Act as a primary responder for well-understood production incidents and participate independently in the on-call rotation.
Assess the impact of incidents, including transaction volume affected, potential revenue impact, and implications for data integrity.
Investigate operational issues using logs, metrics, dashboards, and distributed tracing, then document clear incident updates for stakeholders.
Contribute to postmortems, update runbooks, identify recurring incident patterns, and suggest preventive measures.
Create and maintain dashboards, alerts, and basic service-level indicators for the services you support.
Tune alert thresholds to reduce noise and improve the quality of operational signals.
Participate in system design reviews, with a particular focus on reliability, operability, failure modes, and production readiness.
Implement reliability improvements such as health checks, retries with exponential backoff, circuit breakers, and appropriate monitoring.
Manage cloud resources and contribute to Infrastructure as Code using tools such as Terraform.
Write automation scripts and small internal tools in Bash, Python, or Go to reduce manual toil and improve operational efficiency.
Contribute to CI/CD pipelines and automate routine maintenance tasks such as backup verification, certificate renewal, and log management.
Participate in infrastructure code reviews and help maintain high standards for safe, repeatable changes.
Support security and compliance activities, including PCI DSS controls, ISO 27001 initiatives, security remediation, data classification, and encryption requirements.
Monitor resource utilisation, provide basic capacity forecasts, and implement practical cost optimisation measures such as rightsizing resources and removing unused infrastructure.
Collaborate with development, product, and security teams to improve the resilience and operability of services.
Provide clear handovers, maintain high-quality documentation, and communicate technical topics effectively to both technical and non-technical stakeholders.
Use approved AI tools responsibly to support tasks such as code generation, documentation, and log analysis, while validating outputs and protecting sensitive information.
Your team
Core Infrastructure is responsible for building and operating the foundations that enable Swan’s products to remain reliable as the business grows. We work closely with development and other technical teams to improve system resilience, operational efficiency, and customer experience.
We value ownership, pragmatism, knowledge sharing, and open communication. Engineers are encouraged to challenge ideas constructively, document what they learn, and continuously improve the way we build and operate services. You will work in a supportive environment where reliability is a shared responsibility and where operational excellence is developed through collaboration.
Together alongside Engineering Productivity, our squad constitutes the broader Platform Engineering team.