Production Support Engineer

Solvd, Inc.

Occidente

Presencial

COP 291.290.000 - 388.387.000

Jornada completa

14 días+
Generador de candidaturas

Convierte este puesto en una entrevista — un currículum y una carta de presentación creados pensando en lo que quiere el empleador.

Supera los filtros ATS

Descripción de la vacante

Solvd Inc. is seeking a Production Support Engineer to join a small, high-trust team ensuring the health and reliability of a revenue-critical sales platform. You’ll triage incidents, communicate with partners, and drive post-mortems while avoiding code fixes, focusing on plain-language explanations and stakeholder alignment.

Ideal candidates have strong API/log proficiency, experience with observability tools, SQL, and excellent communication. You’ll work during U.S.

Formación

  • 2+ years troubleshooting apps, servers, or infra.
  • 2+ years providing clear status updates to stakeholders.
  • Experience reading API calls/responses with tools like Postman.
  • Ability to read logs and use observability tools (Splunk, Datadog, Sumo Logic).
  • SQL queries for troubleshooting and ad hoc reporting.
  • Read HTML/JSON and use browser dev tools for investigation.
  • Experience bridging tech and non-technical audiences.
  • On-call readiness with strong time management.
  • Bachelor's degree in related field or equivalent work experience.

Responsabilidades

  • Monitor platform health and triage incidents (outages, degradation).
  • Investigate incidents using logs, API calls, and responses to determine root cause.
  • Own incident management, post-mortems, and corrective actions.
  • Notify stakeholders of critical issues with SLA risk and status updates via email/phone/ticket system.
  • Cross-reference tickets across systems and track defects through lifecycle.
  • Communicate findings clearly to partners and call center teams.
  • Manage incident queue in Jira and prioritize bugs for engineering sprints.
  • Participate in weekly cross-functional meetings with engineering and account management.
  • Provide suggestions for continual process improvement.
  • Join on-call rotations after ramp-up and respond to alerts in defined SLA windows.

Conocimientos

Troubleshooting
API basics
Logging tools
SQL queries
HTML/JSON
Communication
Time management
Customer service
Empathy

Educación

Bachelor's degree

Herramientas

Postman
Splunk
Datadog
Sumo Logic
OpsGenie

Descripción del empleo

Solvd Inc. is a rapidly growing AI-native consulting and technology services firm delivering enterprise transformation across cloud, data, software engineering, and artificial intelligence.

We work with industry-leading organizations to design, build, and operationalize technology solutions that drive measurable business outcomes.

Following the acquisition of Tooploox, a premier AI and product development company, Solvd now offers true end-to-end delivery - from strategic advisory and solution design to custom AI development and enterprise-scale implementation.

Our capability centers combine deep technical expertise, proven delivery methodologies, and sector-specific knowledge to address complex business challenges quickly and effectively.

We are looking for a Production Support Engineer to join a small, high-trust team responsible for the health and reliability of a revenue-critical sales platform. You'll sit between end users - partners and call centers - and engineering, triaging incidents, managing communication, and keeping stakeholders informed and calm when things go wrong.

You'll own the incident management response process end-to-end - from first alert through post-mortem and corrective action follow-up. This is not a developer role. The technical bar is deliberately calibrated: you need to understand how APIs work, read logs, and interpret what you're seeing - not implement fixes. What matters equally is your ability to translate technical issues into plain language and manage expectations across very different audiences.

Longevity and genuine interest in the role matter here. This team values people who want to grow with it, not move through it.

What You'll Do
  • Monitor platform health and triage incoming incidents - distinguishing critical issues (outages, service degradation) from non-critical ones (bugs, defects).
  • Investigate incidents using logging tools - reading API calls, responses, and log data to understand what happened and where.
  • Own the incident management response process, post-mortems, and corrective action follow-up.
  • Notify stakeholders of critical issues proactively - specifying SLA risk and communicating clearly on status via email, phone, or ticket system.
  • Cross-reference tickets across multiple systems and follow defects through the full lifecycle until closure.
  • Communicate clearly with partners and call center teams - translating technical findings into plain language and managing expectations throughout resolution.
  • Manage the incident queue in Jira and prioritize bugs within engineering sprint cycles.
  • Participate in weekly cross-functional meetings with engineering and account/call center management.
  • Provide suggestions for continual improvement of applications and processes.
  • Join on-call rotations after ramp-up - responding to alerts via OpsGenie within defined SLA windows.
Basic Qualifications
  • 2+ years of troubleshooting and resolving issues for applications, servers, or infrastructure environments.
  • 2+ years of providing clear status updates on tasks, issues, and resolutions to stakeholders at multiple levels.
  • Working knowledge of how APIs function - able to read and interpret API calls and responses; experience with Postman or similar API testing tools.
  • Ability to navigate logging and observability tools such as Splunk, Datadog, or Sumo Logic.
  • Experience with SQL queries for troubleshooting and ad hoc reporting.
  • Basic comfort reading HTML and JSON, and using browser developer tools for investigation.
  • Ability to participate in technical bridge calls and follow incidents through to resolution.
  • Exceptional communication skills - able to code-switch between technical and non-technical audiences fluidly; this is the hardest skill to train and the most important one for this role.
  • Strong time management, prioritization, and organizational skills under pressure.
  • Customer service mindset - genuine interest in supporting end users and resolving issues, not just closing tickets.
  • Empathy, humility, and comfort with ambiguity - able to investigate complex issues without a clear playbook.
  • Available during U.S. Eastern business hours (9 AM - 6 PM ET); Eastern timezone strongly preferred for onboarding and on-call coordination.
  • Bachelor's degree in a related field or equivalent work experience.
Preferred Qualifications
  • Experience with AWS - Cloud Practitioner level or above.
  • Familiarity with Git in a team environment.
  • Understanding of infrastructure-as-code concepts - Terraform or similar.
  • Familiarity with OpsGenie or similar alerting platforms.
  • Experience using AI tooling to amplify troubleshooting and investigation workflows.
  • Understanding of engineering deployment lifecycle and release processes.
  • Experience with on-call rotation structures and incident severity frameworks.
  • Prior exposure to partner or call center communication management during live incidents.
  • Experience in a travel, hospitality, or high-volume transactional platform environment.
What To Expect When You Join
  • Comprehensive onboarding documentation and a structured 6-month ramp to full self-sufficiency.
  • On-call rotations begin only when you're ready, with manager backup during early rotations.
  • Active alert window is 8 AM-1 AM Eastern; overnight suppression windows are built in.
  • SEV-1 incidents are rare - roughly once per quarter or less; the majority of the work happens during business hours.
When you join Solvd, you'll…
  • Shape real-world AI-driven projects across key industries, working with clients from startup innovation to enterprise transformation.
  • Be part of a global team with equal opportunities for collaboration across continents and cultures.
  • Thrive in an inclusive environment that prioritizes continuous learning, innovation, and ethical AI standards.

Solvd is an equal opportunity employer.

Consigue la evaluación confidencial y gratuita de tu currículum.

o arrastra y suelta tu archivo aquí

Similar jobs

Puestos de trabajo similares que vale la pena comparar

Production Support Engineer: Incident Response & Communications
Production Support Engineer: Incident Response & Communications

Solvd, Inc. • Occidente

Presencial
COP 291.290.000 - 388.387.000
Senior Full Stack Software Engineer (.Net/C#)
Senior Full Stack Software Engineer (.Net/C#)

Solvd, Inc. • Colombia

Presencial
COP 90.000.000 - 120.000.000
Senior / Lead Full Stack Engineer (TypeScript + React)
Senior / Lead Full Stack Engineer (TypeScript + React)

Solvd, Inc. • Occidente

Presencial
COP 120.000.000 - 180.000.000
Senior Site Reliability Operations Engineer - Finance
Senior Site Reliability Operations Engineer - Finance

Truelogic Software LLC • Bogotá ciudad

Híbrido
COP 388.425.000 - 582.637.000
100% Remote Work
Highly Competitive USD Pay
Paid Time Off
+2
DevOps/Cloud Engineer
DevOps/Cloud Engineer

AgileEngine • Colombia

A distancia
COP 120.000.000 - 210.000.000
Forward Deployed Engineer (TypeScript + AI) - Remote - Colombia
Forward Deployed Engineer (TypeScript + AI) - Remote - Colombia

FullStack • Cartagena de Indias

A distancia
COP 291.290.000 - 420.753.000
100% remote work
Sura health policy
English classes
+3
DevOps Engineer ID89052
DevOps Engineer ID89052

AgileEngine, LLC. • Metropolitana

Presencial
COP 80.000.000 - 120.000.000
Professional growth
Competitive USD-based compensation
Exciting projects
+1
DevOps Engineer ID89052
DevOps Engineer ID89052

AgileEngine, LLC. • Pereira

Híbrido
COP 202.007.000 - 370.345.000
Professional growth
Competitive USD-based compensation
A selection of exciting projects
+1
DevOps Engineer ID89052
DevOps Engineer ID89052

AgileEngine, LLC. • Colombia

Presencial
COP 303.010.000 - 471.349.000
Professional growth
Competitive USD-based compensation
A selection of exciting projects
+1
DevOps Engineer ID89052
DevOps Engineer ID89052

AgileEngine, LLC. • Capital

Híbrido
COP 303.010.000 - 505.016.000
Professional growth
Competitive USD-based compensation
A selection of exciting projects
+1