IT & Service Delivery Engineer
Job Fit Check
Base Career helps you apply smarter for this job.
Key skills for this role
Role Overview
Based at Trakm8's head office in Coleshill, the IT & Service Delivery Engineer is responsible for the availability, performance, security and resilience of Trakm8's telematics and optimisation platforms and of the corporate IT services.
The role spans a mixed environment of corporate office, on-premises data centres and AWS.
Customer platforms process data in real time, alongside the standard business IT services used by employees.
Reliable delivery of both is critical.
Trakm8 runs a small, multi-skilled team covering both IT and Service Delivery.
Between them the team supports over 800 servers, eight production platforms, plus UAT, QA and Dev environments and around 60 users, whilst supporting services 24x7.
Key Skills for This Role
Full Job Posting
Job Purpose
Based at Trakm8's head office in Coleshill, the IT & Service Delivery Engineer is responsible for the availability, performance, security and resilience of Trakm8's telematics and optimisation platforms and of the corporate IT services. The role spans a mixed environment of corporate office, on-premises data centres and AWS. Customer platforms process data in real time, alongside the standard business IT services used by employees. Reliable delivery of both is critical.
Trakm8 runs a small, multi-skilled team covering both IT and Service Delivery. Between them the team supports over 800 servers, eight production platforms, plus UAT, QA and Dev environments and around 60 users, whilst supporting services 24x7.
Service Availability & Incident Management
Monitor the health, capacity and performance of the telematics and optimisation platforms, corporate IT services and supporting infrastructure, taking action to maintain agreed levels of availability and throughput.
Investigate and resolve incidents across Linux and Windows servers, databases, networks, storage, container platforms, end user devices and application deployments, owning each from first report or alert through to root cause and permanent fix.
Carry out structured root cause analysis for recurring or significant issues, implementing corrective and preventative actions.
Communicate incident progress, risks and resolution clearly to technical and non-technical stakeholders.
Infrastructure Engineering & Delivery
Design, build, configure and maintain secure, resilient server infrastructure across data centre, virtualised and AWS environments.
Manage and improve containerised services, networking, storage, databases and load balancing.
Own backup, replication and restore testing, and contribute to capacity planning and disaster recovery, keeping recovery points and recovery times fit for purpose and evidenced.
Work with development, platform and service teams to support reliable application deployment and end-to-end service performance.
Windows, Directory & Endpoint Services
Administer Active Directory, Group Policy, DNS and internal certificate services across the domain.
Own server and workstation patching and endpoint protection coverage, including approvals, safeguard holds, deployment rings.
Administer accounts, access rights and licensing, and maintain inventory, imaging and software deployment tooling.
Respond directly to user requests and access issues, taking ownership through to resolution within the team’s operating model.
Maintain internal IT services including LAN, wireless, VOIP and mobile telephony, and support IT procurement, supplier management and hardware refresh.
Automation, Monitoring & Continuous Improvement
Develop and maintain automation and scripts for repeatable operational tasks.
Maintain effective monitoring, alerting and dashboards across metrics and log platforms, reviewing thresholds and coverage to identify issues early and reduce avoidable incidents.
Build self-healing and auto-remediation where it is safe to do so, for example automated rebalancing of workloads in response to queue lag or resource pressure.
Identify technical debt and resilience gaps and recommend proportionate improvements, making effective use of modern tooling including AI-assisted development and automation.
Security, Documentation & Collaboration
Operate infrastructure with a security-first approach, applying access controls, hardening and vulnerability remediation in line with company policies.
Support logging, monitoring and security tooling, and assist with investigation and remediation of security alerts.
Plan and deliver changes and project work through the RFC and release process, booking into agreed release windows and providing test, rollback and post-implementation evidence.
Produce recurring operational and service reporting covering availability, capacity, patching, backup and security, and maintain accurate asset, licence and configuration records.
Create and maintain clear technical documentation, runbooks and support procedures for new and existing solutions.
Technology Environment
The environment currently includes the following technologies. The list provides context for the role and is not intended to mean that experience in every technology is essential. The expectation is competence in several of these groups and the ability to pick up the rest.
Service delivery platform
Oracle Linux and Windows operating systems across VMware vSphere, Proxmox and AWS
Kubernetes and Docker, using MetalLB and Kong API gateway
MySQL, Cassandra (Scylla), CockroachDB, Redis, Kafka, Elasticsearch, PostgreSQL and RabbitMQ
Ansible and AWX, with scripting using Bash, Python or Perl
TCP/IP networking, DNS, routing, VPNs, BGP, SMTP, HTTP and HTTPS
Palo Alto firewalls; HAProxy, NGINX and dynamic DNS services
Grafana, Prometheus, Graphite (ClickHouse), InfluxDB and Icinga, with alerting into OpsGenie
ELK for application logging and Wazuh for platform host security monitoring
Git, Bitbucket, Jira and Bamboo CI/CD pipelines
iSCSI SAN, storage pools, volumes and snapshots
AWS services including EC2, Route53, RDS, ELB, VPC, Multi-AZ subnets, VPN, Transit Gateway and Customer Gateway
Corporate IT
Windows Server and Windows desktop estates, Active Directory, Group Policy, DNS, DHCP, DFS and file and print services with Linux (Ubuntu)
Microsoft 365 and Entra ID, Intune, including Exchange Online, SharePoint, Teams, multi-factor authentication and Conditional Access
VMware vSphere and vCenter, with Veeam Backup & Replication and object storage for offsite copies
Endpoint management, patching and software deployment tooling, with BitLocker disk encryption
CrowdStrike Falcon endpoint protection and UTMStack SIEM
Palo Alto firewalls managed through Panorama, with GlobalProtect remote access
LAN, wireless, VOIP and mobile telephony across multiple UK sites
Puppet for configuration management and PowerShell for scripting, with Icinga and OpsGenie alerting shared across both estates
Internal certificate services, TLS and PKI, split-horizon DNS, LDAP and Kerberos
MS SQL behind business systems, plus the Atlassian suite, service desk, intranet and digital signage platforms
Essential Skills / Experience / Competency Required
- Strong problem solving, with the ability to pick up an unfamiliar system quickly, retain it, and carry what you learn across into unrelated parts of the estate. Recognising that a fault in one area has the same shape as one you solved somewhere else entirely is worth more here than depth in any single technology.
- Hands-on experience administering Linux in a production, business-critical environment, including fault diagnosis, performance management, patching and security hardening.
- Solid Windows Server administration, including Active Directory and Group Policy.
- Experience administering Microsoft 365 and Entra ID.
- A working understanding of TCP/IP networking, DNS, routing, VPNs and firewall policy: enough to establish whether a fault sits in the application, the host, or a rule upstream.
- A working understanding of Kubernetes and containerised services: enough to operate and diagnose them day to day. Deeper experience is welcome but not essential.
- Scripting and automation in at least one of Bash, Python or PowerShell, and a preference for putting configuration into version control rather than doing things by hand.
- Methodical, evidence-led troubleshooting: working from log, packet and configuration evidence rather than assumption, including faults that do not originate where they first appear.
- Clear written and verbal communication, with the ability to explain technical issues and risks to different audiences.
- A full UK driving licence is required, as occasional travel to data centres or other locations may be required to support IT and platform services.
- Participation in the team's 24x7 on-call rota, with flexible working outside normal hours where reasonably required to protect or restore critical services.
Platform and cloud
Depth in Kubernetes administration, cluster operations and container networking.
AWS, VMware or other virtualisation and cloud platforms.
Databases in production, such as MySQL, PostgreSQL, MS SQL, Cassandra, Kafka or Elasticsearch.
API gateway technologies such as Kong and measuring or reporting API performance.
Telemetry, IoT, high-throughput data platforms or similarly critical real-time services.
Storage, SAN and data centre experience covering physical security, networking and compute.
Corporate IT
Endpoint management, software deployment and patching tooling such as PDQ, WSUS or Intune.
Backup and recovery tooling such as Veeam, including restore testing.
Providing IT support to end users directly, with the confidence to deal with colleagues at any level.
Automation and monitoring
Configuration management with Ansible, Puppet or AWX.
Monitoring, alerting and dashboarding with Icinga, Grafana, Prometheus or similar.
Infrastructure-as-code, CI/CD tooling or automated testing for infrastructure changes.
AI-assisted coding and automation tooling used to accelerate operational and scripting work.
Security and assurance
Firewall policy administration, ideally Palo Alto, and load balancing or reverse proxy work.
Endpoint protection, vulnerability management or SIEM platforms.
System build following NIST standards, hardening, SELinux and server firewalls.
Information security frameworks such as ISO 27001 or Cyber Essentials.
Ways of working
Working within a formal RFC and change window process, including rollback planning and evidence.
Moving live services off unsupported platforms and end-of-life operating systems without an outage.
Disaster recovery, capacity planning and service continuity testing.
Personal Attributes
Calm, methodical and solutions-focused, particularly when responding to live service issues.
Takes ownership and follows actions through to resolution.
Proactive in identifying risks, improvement opportunities and emerging capacity or resilience concerns.
Security-conscious and careful when working with critical systems and customer data.
Collaborative and approachable, with a willingness to share knowledge and support colleagues.
Organised and adaptable, able to balance planned work with changing operational priorities.
Committed to clear documentation, continuous learning and improving how the team works.
Our Values
We’re a close-knit team who care about each other, our customers and making a positive impact, our values drive who we are and the decisions we make:
Caring & Committed – We show up for each other and deliver with dedication
Make a Difference – We create impact that matters
Relentless – We push forward until we succeed
Do the Right Thing – We choose integrity every time
Joy of the Journey – We enjoy the ride and celebrate along the way
Health & Safety
As a Trakm8 Group employee, you have the following H&S responsibilities:
To comply with the Group H&S policy and procedures
To take care of your own health and safety and that of people who may be affected by what you do (or what you do not do)
To co-operate with others on health and safety
To not interfere with, or misuse, anything provided for your health, safety or welfare
To wear/use the correct Protective Equipment at all times
To follow the training you have received at all times
Information Security
As a Trakm8 Group employee, you have the following Information Security responsibilities:
To follow the Group Information Security policies at all times
To follow the Trakm8 Secure Development Policy at all times;
To ensure that passwords under your responsibility are kept confidential
To keep Group or Company information secure and confidential at all times
To inform your manager immediately if you detect, suspect or witness an incident that may be a breach of security
This role requires screening in the below areas
References
Pre-Employment Credit Check
Basic Disclosure
The above job description is meant to describe the general nature and level of work being performed; it is not intended to be construed as an exhaustive list of all responsibilities, duties and skills required for the position. All job descriptions are subject to possible modifications in line with the needs of the business.
About Volarisgroup
Verified company details for this employer are not available yet.
Apply for this job in 1 click
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
More jobs at Volarisgroup
IT & Service Delivery Engineer
, GBR
Bid & Proposal Coordinator
, GBR
Business Development Manager
Northampton, GBR
Integration Manager
, AUS
Project Manager
Abingdon, GBR
QA Analyst
Milton Keynes, GBR
Systems Configuration Analyst
Milton Keynes, GBR
Senior Embedded Linux & Vision AI Engineer
, GBR