platform

package
v0.6.0 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Oct 5, 2026 License: Apache-2.0 Imports: 18 Imported by: 0

Documentation

Overview

Package platform provides model-independent, CSP/GPU-architecture-only overrides for WorkloadRun resources. Override definitions live in YAML templates (overrides/workloadrun.yaml) that reference the same _lib/ fragments used by catalog entries, ensuring a single source of truth for CSP/GPU-specific networking, NCCL configuration, and DRA setup.

Index

Constants

View Source
const (
	AWS        = "aws"
	GCP        = "gcp"
	Azure      = "azure"
	OCI        = "oci"
	OnPrem     = "onprem"
	TogetherAI = "togetherai"
	Mistral    = "mistral"
	Forge      = "forge"
	NScale     = "nscale"
)

Platform name constants. These are the values the controller's platform detection (pkg/controller/workflow_detect.go) can return and the only values the CLI --platform flags accept. Detection returns these constants and the CLIs derive validation, help text, and error messages from Names(), so the two can never drift apart.

View Source
const (
	// JobTemplateWorkloadLabelsPath names workload metadata on a Job or
	// Workflow spec. Callers validating those API surfaces should use it in
	// field-specific errors.
	JobTemplateWorkloadLabelsPath = "spec.jobTemplate.spec.workloadMetadata.labels"
	// WorkloadRunWorkloadLabelsPath names the source field on a WorkloadRun.
	// The generated Workflow has a jobTemplate, but reporting that internal
	// path would point WorkloadRun users at a field their API does not expose.
	WorkloadRunWorkloadLabelsPath = "spec.workloadMetadata.labels"
)

Variables

This section is empty.

Functions

func ApplyGangSchedulerToDependencies added in v0.2.0

func ApplyGangSchedulerToDependencies(deps []nvcrev1alpha1.DependencySpec, gs *nvcrev1alpha1.GangSchedulerSpec) error

ApplyGangSchedulerToDependencies rewrites every TrainingRuntime dependency in place so each of its replicatedJobs runs under the configured gang scheduler. For each replicatedJob it sets schedulerName on the pod spec and the queue label (gangScheduler.queueLabelKey, "kai.scheduler/queue" when unset) on both the job template metadata and the pod template metadata, matching what BuildTorchRuntime and BuildMPIRuntime already emit for a WorkloadRun. The pod-level copy keeps queue assignment from depending on the Trainer/JobSet layer propagating template metadata onto the pods.

For the KAI + MPI shape it also rewrites the JobSet ordering: it deletes the launcher's dependsOn gate together with the JobSet startupPolicy (the two are mutually exclusive) and appends the wait-for-workers init container, because KAI requires every sub-group of the PodGroup to have pods before the group is schedulable — an ordered or gated launcher sub-group is empty until the workers are ready and can never be admitted. The wait therefore lives inside the launcher pod (launcherWaitScript); a dependency between replicated jobs that is not the launcher's node-ready gate keeps its meaning (measured live on KAI v0.16.4; see schedulerNameKAI).

It is a no-op when gs is nil, so a Certification that does not ask for gang scheduling renders byte-identically to before. It is also a no-op when the scheduler name is empty, matching applyGangScheduler: the CRD requires a non-empty schedulerName, but nvcrectl renders straight from a file without consulting the API server, so without this guard a typo would render pod templates pinned to an empty scheduler and carrying a stray queue label.

Callers must invoke this after overrides are resolved. The scheduler name is overwritten unconditionally rather than filled in only when absent, because some catalog entries hardcode schedulerName: default-scheduler and the whole point of the field is to replace it.

func ApplyImageToDependencies added in v0.3.0

func ApplyImageToDependencies(deps []nvcrev1alpha1.DependencySpec, image string) error

ApplyImageToDependencies rewrites every TrainingRuntime dependency in place so the image of the primary workload container, containers[0], of each replicatedJob's pod template is the given image. Init containers whose image exactly equals the primary's pre-override image follow it: the catalog derives those inits from the workload ref (the MPI launcher's fix-ssh-permissions, the training entries' megatron-clone), so they must run whatever image the workload runs, including any source the operator provisioned inside it. Init containers with a distinct image are infrastructure and stay pinned by construction (the GCP tcpxo-daemon init container keeps GCP's RxDM image). Non-init containers beyond index 0 are never touched, whatever their image.

It is a strict no-op when image is empty, so a Certification that never set options.image renders byte-identically to before, matching ApplyGangSchedulerToDependencies's guard. Unlike that helper's ensureMap labels write, it never creates structure: a replicatedJob with no containers list is left alone, byte for byte.

Callers must invoke this after overrides are resolved, at the same point gang scheduling is applied: platform overrides choose images too (the AWS EFA overrides swap the workers to an nccl-tests build), and the whole point of options.image is to replace whatever the catalog landed on. It is a deliberate structural twin of ApplyGangSchedulerToDependencies rather than a shared walker; a third post-resolve rewriter (issue #212 is the candidate) is the refactor trigger.

func ApplyImageToJobTemplate added in v0.3.0

func ApplyImageToJobTemplate(jt *nvcrev1alpha1.JobTemplateSpec, image string)

ApplyImageToJobTemplate points the job template's trainer at the given image (jobTemplate.spec.workload.trainJob.trainer.image), the sibling of ApplyImageToDependencies for the typed half of the workload. It is a no-op when image is empty, and it never creates structure: a template without a trainJob workload or without a trainer has no trainer image to replace and is left alone.

func ApplyResolvedWorkflowTransforms added in v0.5.0

func ApplyResolvedWorkflowTransforms(
	spec *nvcrev1alpha1.WorkflowSpec, t ResolvedWorkflowTransforms,
) error

ApplyResolvedWorkflowTransforms applies every post-resolve transform to a resolved WorkflowSpec in one named stage, then persists the gang-scheduling intent and checks the result.

It exists because a third post-resolve rewrite landed: ADR-077 deliberately left gang scheduling and image replacement as two focused helpers called in sequence at three call sites, and named this the refactor trigger. Those helpers and their tests are unchanged; this composes them.

The order is pinned:

  1. ApplyGangSchedulerToDependencies
  2. ApplyImageToJobTemplate
  3. ApplyImageToDependencies
  4. resolve and write JobTemplate.spec.workloadMetadata

Gang scheduling and image replacement still touch disjoint fields, so only step 4's position matters: writing workload metadata last makes the resolved catalog and override JobSpec the base map, which makes its precedence unambiguous.

It then records the resolved intent on spec.gangScheduler and runs ValidateResolvedJobTemplate, so a caller cannot forget either.

func ApplyWorkloadRunScheduling added in v0.5.0

func ApplyWorkloadRunScheduling(
	spec *nvcrev1alpha1.WorkflowSpec,
	gs *nvcrev1alpha1.GangSchedulerSpec,
	md *nvcrev1alpha1.WorkloadMetadata,
) error

ApplyWorkloadRunScheduling composes the workload-object metadata and records the resolved gang-scheduling intent on a WorkflowSpec that WorkloadRun just built. It is the construction-time half of the same contract ApplyResolvedWorkflowTransforms provides for Certification.

WorkloadRun needs no ApplyGangSchedulerToDependencies call: its runtime dependency is built with the scheduler and queue already in place by BuildTorchRuntime, BuildMPIRuntime and BuildExecRuntime, unlike a catalog runtime that ADR-076 has to rewrite after overrides resolve.

This merge is provisional. WorkloadRun's structural overrides run later and may change either metadata level, so the persisted intent is what makes the Workflow controller's post-override check authoritative.

func BaseNCCLEnvVars

func BaseNCCLEnvVars(enableMNNVL bool) []corev1.EnvVar

BaseNCCLEnvVars returns model-independent NCCL environment variables that are safe to auto-inject for any workload. These are common across all training models and NCCL tests.

Platform-specific NCCL vars (FastRak, IB settings) are injected via overrides, not here.

func BuildExecRuntime

func BuildExecRuntime(cfg RuntimeConfig) nvcrev1alpha1.DependencySpec

BuildExecRuntime creates a TrainingRuntime dependency for arbitrary command execution. Uses a simple single-replicatedJob layout with torch mlPolicy.

func BuildMPIRuntime

func BuildMPIRuntime(cfg RuntimeConfig) nvcrev1alpha1.DependencySpec

BuildMPIRuntime creates TrainingRuntime dependencies for MPI-based workloads. Generates: - A runtime with MPI mlPolicy and launcher+node replicatedJobs - Worker nodes with sshd, IPC_LOCK, readiness probe, and cfg.Env - Launcher with mpirun, SSH key setup, and cfg.Env

The worker's openssh-server install is guarded by `test -x /usr/sbin/sshd`, so images that already ship sshd start without any package-manager egress (issue #319: air-gapped and restricted-egress clusters). The guard tests the exact path the chain execs (/usr/sbin/sshd); a PATH lookup could disagree with it in either direction.

cfg.Env goes on both containers as container-level env, the same way BuildTorchRuntime emits it (issue #68: it used to be dropped here, so spec.env behaved differently between the two frameworks). On the launcher the variables reach mpirun directly; on the worker they reach the sshd process, but sshd gives each SSH session a fresh environment, so ranks launched through it may not inherit them. A variable that must reach the ranks should be passed as "-x NAME=value" in mpiArgs, which forwards it through mpirun itself.

func BuildTorchRuntime

func BuildTorchRuntime(cfg RuntimeConfig) nvcrev1alpha1.DependencySpec

BuildTorchRuntime creates a TrainingRuntime dependency for PyTorch distributed training. Generates a runtime with torch mlPolicy and a single "node" replicatedJob.

func MergeEnvVars

func MergeEnvVars(base, user []corev1.EnvVar) []corev1.EnvVar

MergeEnvVars merges base env vars with user-provided env vars. User-provided values take precedence (override by name).

func Names

func Names() []string

Names returns every valid platform name, in the order used in flag help and error messages.

func NamesList

func NamesList() string

NamesList returns the valid platform names joined with ", " for flag help and error messages.

func PreserveSchedulingFields added in v0.5.0

func PreserveSchedulingFields(
	original, renamed []byte, gs *nvcrev1alpha1.GangSchedulerSpec,
) ([]byte, error)

PreserveSchedulingFields copies the gang-scheduling fields ADR-076 writes from original into renamed, and returns the corrected document.

It exists because per-job dependency renaming rewrites every quoted occurrence of a dependency name, which also catches a queue label value that happens to equal one — a queue named after the runtime it serves. The queue a user configured is not a reference and must survive renaming intact.

Crucially this preserves rather than repairs. Unlike ApplyGangSchedulerToDependencies it asserts nothing and inserts nothing: a field absent from original stays absent, and a value the operator overrode to something else is carried through unchanged. An override that redirects the queue therefore survives renaming and is rejected by validation, which is the reject-without-repair contract. Re-applying the configured intent here instead would silently overwrite that override and let the Job run in a queue the operator did not choose.

It is a no-op when gs configures nothing, since then there is no queue label key to identify.

func ResolvedGangScheduler added in v0.5.0

ResolvedGangScheduler returns a copy of gs with queue and queueLabelKey filled in with the defaults the write paths already apply, so the value can be persisted on a Workflow and later compared without re-deriving them. It returns nil when gs configures nothing, matching the guard in ApplyGangSchedulerToDependencies: the CRD requires a non-empty schedulerName, but nvcrectl renders straight from a file without consulting the API server.

func ValidateFlag

func ValidateFlag(name string) error

ValidateFlag validates a --platform flag value. An empty value is allowed (platform not specified). The error text derives the valid names from Names() so the message always matches the actual set.

func ValidateResolvedGangScheduling added in v0.5.0

func ValidateResolvedGangScheduling(
	js *nvcrev1alpha1.JobSpec,
	deps []nvcrev1alpha1.DependencySpec,
	gs *nvcrev1alpha1.GangSchedulerSpec,
	workloadLabelsPath string,
) error

ValidateResolvedGangScheduling checks that the effective manifests NVCRE is about to submit still place the workload in the queue its owner configured, and restores the one label NVCRE itself owns.

The asymmetry between the two halves is deliberate:

  • The workload-object queue label is NVCRE's to write. It follows the insert/same/conflict rule, so an override that removed it gets it back and an override that changed it fails.
  • Runtime queue labels and pod scheduler names are assertions over dependency payloads the operator may have redirected. A missing or conflicting runtime value fails and is never repaired, because silently rewriting someone's runtime override would hide the disagreement rather than surface it.

js is the working copy used to create children, not the stored Workflow, so restoring the label here does not touch the Workflow's immutable metadata.

A nil gs is a no-op. Callers wanting the checks that apply whatever the scheduling configuration should use ValidateResolvedJobTemplate instead.

gs is re-resolved rather than trusted. NVCRE persists it already resolved, but queue and queueLabelKey are optional on the CRD, so a hand-authored Workflow can name only a schedulerName. Defaulting here keeps that case checking the same queue the write paths would have used instead of an empty key. Re-resolving an already-resolved value changes nothing.

func ValidateResolvedJobTemplate added in v0.5.0

func ValidateResolvedJobTemplate(
	js *nvcrev1alpha1.JobSpec,
	deps []nvcrev1alpha1.DependencySpec,
	gs *nvcrev1alpha1.GangSchedulerSpec,
	workloadLabelsPath string,
) error

ValidateResolvedJobTemplate checks a resolved Job spec against everything NVCRE promises about it, and is the entry point every post-override call site should use.

Workload metadata is validated whatever the scheduling configuration. Admission covers a Job or Workflow that reaches an API server, but offline `nvcrectl render` never does, and an override can introduce a reserved key or an oversized map after the owning resource was admitted. Gating that check behind gangScheduler, as ValidateResolvedGangScheduling alone does, would let resolved offline output carry labels the cluster will reject.

The gang-scheduling contract is then checked when an intent is persisted.

Types

type OverrideConfig

type OverrideConfig struct {
	EntryName   string
	NodesPerJob int32
	GpusPerNode int32
	MlnxPerNode int32

	// NicResourceName is the extended resource name of the RDMA NIC devices
	// for the on-prem GB200/GB300 override. Empty means the override omits
	// the NIC resource block. The per-container count comes from MlnxPerNode.
	NicResourceName string

	EnableMNNVL   bool
	FrameworkType string

	// UserEnv overrides matching names in platform trainer.env patches so
	// WorkloadRun values retain precedence when Kubeflow Trainer applies the
	// patch. User-only variables stay in the runtime container environment.
	UserEnv []corev1.EnvVar `json:"-" yaml:"-"`
}

OverrideConfig holds template data for rendering platform overrides. Field names match the catalog template data shape so _lib/ fragments can be rendered directly without mapping.

type ResolvedWorkflowTransforms added in v0.5.0

type ResolvedWorkflowTransforms struct {
	// GangScheduler is the owner's gang-scheduling intent, before defaulting.
	GangScheduler *nvcrev1alpha1.GangSchedulerSpec

	// Image replaces the workload container image. Empty means no override.
	Image string

	// WorkloadLabels are the workload-object labels the owner resolved from
	// its own API surface — for Certification, global labels already merged
	// with per-category labels.
	WorkloadLabels map[string]string
}

ResolvedWorkflowTransforms carries the intents applied to a WorkflowSpec after catalog and platform overrides have resolved.

type RuntimeConfig

type RuntimeConfig struct {
	// EntryName is the WorkloadRun name (used as prefix for resource names).
	// Matches catalog template variable .EntryName.
	EntryName string
	// Image is the container image.
	Image string
	// NodesPerJob is the number of nodes.
	NodesPerJob int32
	// GpusPerNode is the number of GPUs per node.
	GpusPerNode int32
	// Env is the merged env vars (base NCCL + user).
	Env []corev1.EnvVar
	// Volumes are additional volumes.
	Volumes []corev1.Volume
	// VolumeMounts are additional volume mounts.
	VolumeMounts []corev1.VolumeMount
	// InitContainers are user-provided init containers.
	InitContainers []corev1.Container
	// Resources overrides GPU/memory/CPU resources.
	Resources *corev1.ResourceRequirements
	// ImagePullSecrets for container pull.
	ImagePullSecrets []corev1.LocalObjectReference
	// GangSchedulerName is the scheduler name to inject into pod specs (e.g. "kai-scheduler").
	// Empty means no gang scheduler is configured.
	GangSchedulerName string
	// GangSchedulerQueue is the queue label value for the gang scheduler.
	// Defaults to "default-queue" when GangSchedulerName is set and Queue is empty.
	GangSchedulerQueue string
	// GangSchedulerQueueLabelKey is the label key the queue value is written
	// under. Defaults to "kai.scheduler/queue" when GangSchedulerName is set
	// and the key is empty.
	GangSchedulerQueueLabelKey string
}

RuntimeConfig holds all parameters needed to build a TrainingRuntime dependency.

type WorkloadRunOverride

type WorkloadRunOverride struct {
	nvcrev1alpha1.OverrideSpec

	// PreCommand contains shell lines prepended to the trainer command.
	// Applied by the WorkloadRun controller; baked into trainer.command/args.
	PreCommand []string

	// MPIArgs contains mpirun arguments prepended to the MPI launcher command.
	// Applied by the WorkloadRun controller; baked into trainer.args.
	MPIArgs []string
}

WorkloadRunOverride extends OverrideSpec with fields that are consumed by the WorkloadRun controller at build time and never stored in the Kubernetes API. Keeping them out of OverrideSpec avoids CRD schema changes.

func BuildOverrides

func BuildOverrides(cfg OverrideConfig) []WorkloadRunOverride

BuildOverrides renders the platform override templates and returns the resulting WorkloadRunOverride list.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL