03 ago
|
Ultimate.ai
|
Perú
The Role
Join us as a Site Reliability Engineer (SRE) and embark on an exciting journey of ensuring reliability, resiliency, and innovation in our information systems and ecosystems. As an SRE at Kyndryl, you'll be at the forefront of driving continuous improvement and delivering exceptional service to our customers.
You'll analyze business needs, tackle complex problems, and provide strategic advice and designs. Your responsibilities span every stage of the software lifecycle, from building and testing to deploying changes and maintaining robust systems.
We are looking for a visionary who can think strategically and help shape the future of our services. Your expertise in building trusted relationships with customers and partnering for success will be instrumental in driving our growth.
You will work on end‑to‑end services, spanning customer sites and platforms. Collaboration and proactivity are key as you work with a talented team, taking ownership and continuously seeking innovative solutions.
With an unwavering focus on quality, robustness, and security, you will implement cutting‑edge tools that enhance operations, improve reliability, and gather valuable feedback on our platforms. Your ability to identify and mitigate operational issues will play a crucial role in delivering seamless experiences to our customers.
We want you to shape the future of reliability engineering in a collaborative, entrepreneurial environment. As a Site Reliability Engineer at Kyndryl, you will have opportunities to work on integral projects and collaborate with colleagues worldwide.
Who You Are
You are customer‑focused, growth‑oriented, and inclusive. You possess the required experience and a growth mindset, driving your own personal and professional development.
Required Skills and Experience
- 10+ years of experience in operational management, including incident management and escalations
- Experience with design and implementation of application monitoring to ensure reliability and performance meets or exceeds business goals
- Experience implementing strategies to cap operations load and handle overflow using appropriate tooling and metrics; defining service level indicators and objectives in collaboration with stakeholders, business, development, DevSecOps, and operations teams
- Solution and design experience in an enterprise environment: Windows and Linux servers (RHEL preferred), UNIX (AIX, Solaris), storage, and hyperscaler cloud (AWS, Azure, Google Cloud Platform); public cloud platforms such as AWS, OpenShift, Azure, or GCP
- Experience working with data formats and scripting languages JSON, YAML, Bash, and/or PowerShell
Preferred Skills and Experience
- BS degree in Computer Science, Engineering, or another highly technical, scientific discipline
- Expertise with Ansible, Terraform, and Python
- Experience with distributed technologies and dynamic resource management frameworks such as Kubernetes
- Expertise in leveraging open‑source tooling such as Prometheus, Grafana, or Loki
What You Can Expect
Our career path offers a dynamic, hybrid‑friendly culture supporting your well‑being and growth. You’ll work on impactful projects that sharpen skills and fuel growth.
We champion your journey with powerful tools to chart your career path, personalized development goals, and continuous feedback. You’ll develop in‑demand skills through certifications with Microsoft, Google, and Amazon, coaching, and hands‑on experiences.
Our culture values empathy, restless learning, and shared success, providing an environment where belonging and belonging drives engagement.
Ready to make an impact? Join us and help shape what’s next.
📌 Site Reliability Engineer (Perú)
🏢 Ultimate.ai
📍 Perú