{bc}
workday

IT & Service Delivery Engineer

Volarisgroup
England, GBR
Full-time
Mid
Onsite
Discovered Today
Oracle LinuxWindows ServerVMware vSphereProxmoxAWSKubernetes
Free

Job Fit Check

Base Career helps you apply smarter for this job.

?%
Ready to Scan

Key skills for this role

Oracle LinuxWindows ServerVMware vSphere
Smart Apply

Full Job Posting

Job Purpose

Based at Trakm8's head office in Coleshill, the IT & Service Delivery Engineer is responsible for the availability, performance, security and resilience of Trakm8's telematics and optimisation platforms and of the corporate IT services. The role spans a mixed environment of corporate office, on-premises data centres and AWS. Customer platforms process data in real time, alongside the standard business IT services used by employees. Reliable delivery of both is critical.

Trakm8 runs a small, multi-skilled team covering both IT and Service Delivery. Between them the team supports over 800 servers, eight production platforms, plus UAT, QA and Dev environments and around 60 users, whilst supporting services 24x7.

Service Availability & Incident Management

Monitor the health, capacity and performance of the telematics and optimisation platforms, corporate IT services and supporting infrastructure, taking action to maintain agreed levels of availability and throughput.

Investigate and resolve incidents across Linux and Windows servers, databases, networks, storage, container platforms, end user devices and application deployments, owning each from first report or alert through to root cause and permanent fix.

Carry out structured root cause analysis for recurring or significant issues, implementing corrective and preventative actions.

Communicate incident progress, risks and resolution clearly to technical and non-technical stakeholders.

Infrastructure Engineering & Delivery

Design, build, configure and maintain secure, resilient server infrastructure across data centre, virtualised and AWS environments.

Manage and improve containerised services, networking, storage, databases and load balancing.

Own backup, replication and restore testing, and contribute to capacity planning and disaster recovery, keeping recovery points and recovery times fit for purpose and evidenced.

Work with development, platform and service teams to support reliable application deployment and end-to-end service performance.

Windows, Directory & Endpoint Services

Administer Active Directory, Group Policy, DNS and internal certificate services across the domain.

Own server and workstation patching and endpoint protection coverage, including approvals, safeguard holds, deployment rings.

Administer accounts, access rights and licensing, and maintain inventory, imaging and software deployment tooling.

Respond directly to user requests and access issues, taking ownership through to resolution within the team’s operating model.

Maintain internal IT services including LAN, wireless, VOIP and mobile telephony, and support IT procurement, supplier management and hardware refresh.

Automation, Monitoring & Continuous Improvement

Develop and maintain automation and scripts for repeatable operational tasks.

Maintain effective monitoring, alerting and dashboards across metrics and log platforms, reviewing thresholds and coverage to identify issues early and reduce avoidable incidents.

Build self-healing and auto-remediation where it is safe to do so, for example automated rebalancing of workloads in response to queue lag or resource pressure.

Identify technical debt and resilience gaps and recommend proportionate improvements, making effective use of modern tooling including AI-assisted development and automation.

Security, Documentation & Collaboration

Operate infrastructure with a security-first approach, applying access controls, hardening and vulnerability remediation in line with company policies.

Support logging, monitoring and security tooling, and assist with investigation and remediation of security alerts.

Plan and deliver changes and project work through the RFC and release process, booking into agreed release windows and providing test, rollback and post-implementation evidence.

Produce recurring operational and service reporting covering availability, capacity, patching, backup and security, and maintain accurate asset, licence and configuration records.

Create and maintain clear technical documentation, runbooks and support procedures for new and existing solutions.

Technology Environment

The environment currently includes the following technologies. The list provides context for the role and is not intended to mean that experience in every technology is essential. The expectation is competence in several of these groups and the ability to pick up the rest.

Service delivery platform

Oracle Linux and Windows operating systems across VMware vSphere, Proxmox and AWS

Kubernetes and Docker, using MetalLB and Kong API gateway

MySQL, Cassandra (Scylla), CockroachDB, Redis, Kafka, Elasticsearch, PostgreSQL and RabbitMQ

Ansible and AWX, with scripting using Bash, Python or Perl

TCP/IP networking, DNS, routing, VPNs, BGP, SMTP, HTTP and HTTPS

Palo Alto firewalls; HAProxy, NGINX and dynamic DNS services

Grafana, Prometheus, Graphite (ClickHouse), InfluxDB and Icinga, with alerting into OpsGenie

ELK for application logging and Wazuh for platform host security monitoring

Git, Bitbucket, Jira and Bamboo CI/CD pipelines

iSCSI SAN, storage pools, volumes and snapshots

AWS services including EC2, Route53, RDS, ELB, VPC, Multi-AZ subnets, VPN, Transit Gateway and Customer Gateway

Corporate IT

Windows Server and Windows desktop estates, Active Directory, Group Policy, DNS, DHCP, DFS and file and print services with Linux (Ubuntu)

Microsoft 365 and Entra ID, Intune, including Exchange Online, SharePoint, Teams, multi-factor authentication and Conditional Access

VMware vSphere and vCenter, with Veeam Backup & Replication and object storage for offsite copies

Endpoint management, patching and software deployment tooling, with BitLocker disk encryption

CrowdStrike Falcon endpoint protection and UTMStack SIEM

Palo Alto firewalls managed through Panorama, with GlobalProtect remote access

LAN, wireless, VOIP and mobile telephony across multiple UK sites

Puppet for configuration management and PowerShell for scripting, with Icinga and OpsGenie alerting shared across both estates

Internal certificate services, TLS and PKI, split-horizon DNS, LDAP and Kerberos

MS SQL behind business systems, plus the Atlassian suite, service desk, intranet and digital signage platforms

Essential Skills / Experience / Competency Required

  • Strong problem solving, with the ability to pick up an unfamiliar system quickly, retain it, and carry what you learn across into unrelated parts of the estate. Recognising that a fault in one area has the same shape as one you solved somewhere else entirely is worth more here than depth in any single technology.
  • Hands-on experience administering Linux in a production, business-critical environment, including fault diagnosis, performance management, patching and security hardening.
  • Solid Windows Server administration, including Active Directory and Group Policy.
  • Experience administering Microsoft 365 and Entra ID.
  • A working understanding of TCP/IP networking, DNS, routing, VPNs and firewall policy: enough to establish whether a fault sits in the application, the host, or a rule upstream.
  • A working understanding of Kubernetes and containerised services: enough to operate and diagnose them day to day. Deeper experience is welcome but not essential.
  • Scripting and automation in at least one of Bash, Python or PowerShell, and a preference for putting configuration into version control rather than doing things by hand.
  • Methodical, evidence-led troubleshooting: working from log, packet and configuration evidence rather than assumption, including faults that do not originate where they first appear.
  • Clear written and verbal communication, with the ability to explain technical issues and risks to different audiences.
  • A full UK driving licence is required, as occasional travel to data centres or other locations may be required to support IT and platform services.
  • Participation in the team's 24x7 on-call rota, with flexible working outside normal hours where reasonably required to protect or restore critical services.

Platform and cloud

Depth in Kubernetes administration, cluster operations and container networking.

AWS, VMware or other virtualisation and cloud platforms.

Databases in production, such as MySQL, PostgreSQL, MS SQL, Cassandra, Kafka or Elasticsearch.

API gateway technologies such as Kong and measuring or reporting API performance.

Telemetry, IoT, high-throughput data platforms or similarly critical real-time services.

Storage, SAN and data centre experience covering physical security, networking and compute.

Corporate IT

Endpoint management, software deployment and patching tooling such as PDQ, WSUS or Intune.

Backup and recovery tooling such as Veeam, including restore testing.

Providing IT support to end users directly, with the confidence to deal with colleagues at any level.

Automation and monitoring

Configuration management with Ansible, Puppet or AWX.

Monitoring, alerting and dashboarding with Icinga, Grafana, Prometheus or similar.

Infrastructure-as-code, CI/CD tooling or automated testing for infrastructure changes.

AI-assisted coding and automation tooling used to accelerate operational and scripting work.

Security and assurance

Firewall policy administration, ideally Palo Alto, and load balancing or reverse proxy work.

Endpoint protection, vulnerability management or SIEM platforms.

System build following NIST standards, hardening, SELinux and server firewalls.

Information security frameworks such as ISO 27001 or Cyber Essentials.

Ways of working

Working within a formal RFC and change window process, including rollback planning and evidence.

Moving live services off unsupported platforms and end-of-life operating systems without an outage.

Disaster recovery, capacity planning and service continuity testing.

Personal Attributes

Calm, methodical and solutions-focused, particularly when responding to live service issues.

Takes ownership and follows actions through to resolution.

Proactive in identifying risks, improvement opportunities and emerging capacity or resilience concerns.

Security-conscious and careful when working with critical systems and customer data.

Collaborative and approachable, with a willingness to share knowledge and support colleagues.

Organised and adaptable, able to balance planned work with changing operational priorities.

Committed to clear documentation, continuous learning and improving how the team works.

Our Values

We’re a close-knit team who care about each other, our customers and making a positive impact, our values drive who we are and the decisions we make:

Caring & Committed – We show up for each other and deliver with dedication

Make a Difference – We create impact that matters

Relentless – We push forward until we succeed

Do the Right Thing – We choose integrity every time

Joy of the Journey – We enjoy the ride and celebrate along the way

Health & Safety

As a Trakm8 Group employee, you have the following H&S responsibilities:

To comply with the Group H&S policy and procedures

To take care of your own health and safety and that of people who may be affected by what you do (or what you do not do)

To co-operate with others on health and safety

To not interfere with, or misuse, anything provided for your health, safety or welfare

To wear/use the correct Protective Equipment at all times

To follow the training you have received at all times

Information Security

As a Trakm8 Group employee, you have the following Information Security responsibilities:

To follow the Group Information Security policies at all times

To follow the Trakm8 Secure Development Policy at all times;

To ensure that passwords under your responsibility are kept confidential

To keep Group or Company information secure and confidential at all times

To inform your manager immediately if you detect, suspect or witness an incident that may be a breach of security

This role requires screening in the below areas

References

Pre-Employment Credit Check

Basic Disclosure

The above job description is meant to describe the general nature and level of work being performed; it is not intended to be construed as an exhaustive list of all responsibilities, duties and skills required for the position. All job descriptions are subject to possible modifications in line with the needs of the business.

Apply for this job in 1 click

Skip the repetitive application forms

Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.

Sarah M.James T.Maya R.

Trusted by over 500,000 job seekers on Base Career

Start Free Today

More from this employer

More jobs at Volarisgroup