01
Observe
Inputs
Scheduler queue state, allocations, GPU metrics (DCGM-class), job metadata
Outputs
View of requested vs useful capacity, idle hold, fragmentation, wait pressure
Product
Ainek is being built as the capacity layer between your workloads and your existing scheduler — to observe GPU use, apply team policy, pack and reclaim waste, and explain consumption. Private preview. Not generally available.
For platform teams who want to shape this with us.
02 / Architecture
The design: Ainek decides and records capacity policy. Kubernetes or Slurm still places work on nodes. Device plugins and GPU operators remain yours.
API, policy engine, capacity controller, and usage store — quotas, priorities, packing/reclaim decisions, and attribution.
Cluster nodes, GPU operator / device plugin, trainers, and the scheduler you already run.
Model weights and datasets never need to flow through Ainek. We work from job metadata, allocations, and GPU metrics.
Design partners start in observe-only or suggest mode. Enforce comes later — when policy is trusted. You choose when anything can block, reclaim, or preempt.
Intended default: fail-open for scheduling so your cluster continues under existing Kubernetes or Slurm rules. Exact enforce-mode behavior will be documented with partners before any enforcement is enabled.
03 / Control loop
The loop we're building. Each stage has clear inputs and outputs so platform engineers can reason about behavior — and challenge the design with us.
01
Inputs
Scheduler queue state, allocations, GPU metrics (DCGM-class), job metadata
Outputs
View of requested vs useful capacity, idle hold, fragmentation, wait pressure
02
Inputs
Team quotas, priority classes, fair-share policy, incoming job requests
Outputs
Admit, defer, or reshape requests before scarce GPUs are committed
03
Inputs
Placement constraints, fragmentation, idle reservations, preemption policy
Outputs
Tighter packing and reclaim of unused hold so waiting work can run
04
Inputs
Runtime usage, team identity, policy decisions
Outputs
Usage records and exports for chargeback, review, and capacity planning
04 / Capacity flow
How the system is meant to behave: Team A training, Team B batch inference, an on-call priority job. Ainek in the decision path — the scheduler still places the work.
01
Jobs enter through your normal path — Kubernetes or Slurm — with team and priority context.
02
Quotas, fair-share, and priority evaluated against current fleet pressure.
03
Allowed work proceeds to the scheduler for node placement under your existing rules.
04
GPUs execute. Useful utilization versus reserved hold stays visible.
05
If policy allows, idle capacity is reclaimed or lower-priority work yields for the on-call job.
06
Usage is recorded to the responsible team for review and cost accountability.
05 / Integrations
We're integrating at the scheduling and telemetry boundaries. No rip-and-replace of your training stack. Surfaces below are the target design — refined with design partners.
From you
Cluster API, GPU operator / device plugin, namespaces, labels and taints
What we add
Scheduling hooks and policy signals for shared capacity across teams
From you
Partitions, queues, accounts, and existing fair-share conventions
What we add
Capacity policy and reclaim aligned with how research and product share the pool
From you
Team and identity mapping from your directory or platform source of truth
What we add
Quotas, priorities, and attribution bound to those teams
From you
Existing metrics and alerting pipeline
What we add
Capacity and usage signals you can scrape or export alongside cluster metrics
06 / Deploy model
How we intend to ship with design partners: components in or adjacent to your cluster, progressive risk, you keep control of the fleet.
In-cluster controllers, or a hybrid control plane with agents and metrics collectors beside your nodes — chosen to match your network and ops model.
Access to the cluster API and GPU / scheduler metrics required for observe and policy. Least privilege; scoped per environment.
Observe-only → policy suggest → enforce. Each stage has an exit. Nothing is assumed generally available.
07 / Metrics
What we care about measuring with partners. No fabricated product screenshots — evaluation is a shared review on real fleets.
Week 0
Observe-only. Capture current utilization, wait, and contention patterns.
Suggest
Introduce quotas and priorities in suggest mode. Compare against baseline.
Later
Only when ready: packing and reclaim where safe. Measure wait and useful capacity.
Ongoing
Attribution, waste, and whether software recovered capacity before buying more GPUs.
08 / Boundaries
We're early. Scope discipline matters more than a feature list.
Request early access. Tell us your cluster topology and scheduler — if you're a fit as a design partner, we'll follow up.