Base Career helps you apply smarter for this job.
Key skills for this role
Runware is building a high-performance, full-stack AI media-creation platform — empowering developers and companies to generate any type of media instantly. As we scale fast and integrate increasingly complex models, we need stronger visibility, analytics, and monitoring across the whole platform stack.
We’re looking for a Data Expert (Analytics + Monitoring + Observability) to help us better understand, measure, and optimize how the Runware platform performs at scale — internally and for our clients.
Your main goal is to give Runware full visibility over:
End-to-end inference performance
Integration usage and model activity
Errors, delays, bottlenecks, regressions
Internal and client-facing analytics dashboards
Health and performance of production pipelines
You will provide the data insights that allow engineering, ML, backend, DevOps, and leadership to make informed decisions — and to continuously improve performance and reliability.
Build and maintain E2E inference time tracking (global and per-model).
Monitor how implementation changes impact total request latency.
Detect regressions introduced by suboptimal code paths.
Provide automated alerts & historical trends.
Build dashboards for internal use (engineering, product, leadership).
Skip the repetitive application forms
Install the Base Career Chrome Extension and autofill job applications across major job boards with your profile.
Trusted by over 500,000 job seekers on Base Career
More from this employer
, GBR
, GBR
, GBR
, GBR
, GBR
, GBR
, GBR
, GBR
Provide client-facing usage dashboards (requests, errors, success rate, performance).
Support clients who need visibility to debug their integrations.
Track model-level usage, API endpoints usage, adoption metrics, etc.
Implement metrics, logs, and traces that help the entire platform scale smoothly.
Work closely with DevOps & backend teams to improve system observability.
Provide insights that guide infra decisions (GPU allocation, autoscaling, caching, batching, etc.).
Select and maintain tooling (e.g., Prometheus/Grafana, Datadog, OpenTelemetry, ELK, BigQuery, etc.).
Ensure data pipelines are reliable, accessible, and always up-to-date.
Build simple, easy-to-read dashboards for both technical and non-technical teams.
Generative AI API provider helping developers and businesses create image and visual media.
Visit company websiteJobs and hiring trendsFull-time
Mid
Remote
Apply faster on company sites with our extension.