Dark GPU server hall with a row of 8-GPU server trays receding into the distance

Australian AI Infrastructure

From the GPU to the model endpoint.

One partner for your AI infrastructure.

Plan, build and operate production AI infrastructure — or access GPU capacity from Peregrine's Australian compute fleet.

Use one service or all four. We can work on Peregrine capacity, infrastructure you own, or alongside the cloud environment you already use.

APPLICATION / AI PRODUCTMODEL ENDPOINTSAI PLATFORMGPU INFRASTRUCTURECAPACITY + PLANNINGPEREGRINE
NVIDIA Inception Program badge

Peregrine Compute is a member of the NVIDIA Inception Program.

One partner across the AI infrastructure stack

One partner from the GPU to the model endpoint.

Use one service or all four. Each works on Peregrine capacity, on hardware you own, or alongside the cloud infrastructure you already use.

  1. MODEL ENDPOINTS

    • Inference APIs
    • Fine-tuning
    • MLOps
    • Model serving
    OPERATE
  2. AI PLATFORM

    • Kubernetes
    • Slurm
    • Scheduling
    • Quotas
    • Monitoring
    • Workload management
    OPERATEBUILD
  3. GPU INFRASTRUCTURE

    • GPU servers
    • Networking
    • Storage
    • Cluster architecture
    • Drivers
    • High-speed interconnect
    BUILDCAPACITY
  4. PLANNING

    • Workload sizing
    • Accelerator selection
    • Capacity forecasting
    • Cost modelling
    • Architecture
    ASSESS

Customers can engage at any layer. GPU capacity is only one part of production AI infrastructure.

Four services · one partner

Assess, build, operate and access AI compute infrastructure.

Peregrine Compute helps organisations assess, build, operate and access AI compute infrastructure — from infrastructure planning to production model endpoints.

01ASSESS

Compute Blueprint

Know what to build before you buy compute.

Peregrine analyses your workloads, performance requirements, utilisation forecasts and infrastructure constraints to design an appropriate AI compute architecture.

Commercial model

Fixed fee

Engagement

Typical engagement: approximately 2 weeks.

Deliverable

COMPUTE BLUEPRINT

Request a Compute Blueprint
Architecture drawings and a rack elevation diagram on a steel table in front of a server rack

Blueprint process

  1. Workload
  2. GPU Requirement
  3. Infrastructure
  4. Cost Model
  5. Recommended Architecture

Included

  • Workload sizing
  • GPU or accelerator selection
  • Training and inference requirements
  • Storage and networking requirements
  • Buy vs rent vs cloud comparison
  • Capacity forecasting
  • Data residency requirements
  • Security and compliance mapping
  • Estimated infrastructure economics
  • Scaling strategy
02BUILD

Infrastructure You Own

Production GPU infrastructure, designed and commissioned for you.

Peregrine can design, source, configure and commission GPU infrastructure that your organisation owns.

YOUR PREMISESAUSTRALIAN DATA CENTRECOLOCATIONHYBRID

Commercial model

Fixed-price project

At completion

HANDOVER or PEREGRINE OPERATE

Build My GPU Infrastructure
Open server rack with GPU servers mid-installation and structured cabling

Delivery process

  1. Design
  2. Source
  3. Install
  4. Configure
  5. Test
  6. Handover / Operate

Included

  • GPU server specification
  • Accelerator selection
  • Server sourcing
  • Cluster architecture
  • High-speed networking
  • Storage architecture
  • Rack and power planning
  • NVIDIA drivers
  • CUDA environment
  • Container runtime
  • Kubernetes or Slurm
  • Monitoring
  • Testing
  • Commissioning
  • Documentation
  • Handover
03OPERATE

Managed AI Platform

We operate the AI infrastructure. Your team builds the AI.

Peregrine operates the platform layer between GPU infrastructure and your AI applications.

Operates on: Peregrine GPU infrastructure · Customer-owned · Colocated · Public cloud · Hybrid

ANY HARDWARE.

Operate Peregrine infrastructure, customer-owned infrastructure or hybrid environments through a consistent AI platform.

Commercial model

Monthly managed service

Discuss Managed AI Infrastructure
Glass wall reflecting monitoring dashboards with GPU racks behind

Architecture

  1. AI TEAM
  2. API / NOTEBOOK / DEV ENVIRONMENT
  3. PEREGRINE MANAGED AI PLATFORM
  4. KUBERNETES / SLURM
  5. GPU INFRASTRUCTURE
  6. PEREGRINE / CUSTOMER / CLOUD

Capabilities

  • Kubernetes
  • Slurm
  • GPU scheduling
  • Resource quotas
  • Multi-tenancy
  • User management
  • Monitoring
  • Logging
  • GPU utilisation monitoring
  • Container environments
  • Model serving
  • OpenAI-compatible APIs*
  • Inference endpoints
  • Fine-tuning pipelines
  • MLOps
  • Updates
  • Patch management
  • Capacity management
  • Workload orchestration
  • Storage integration
  • Observability

* Model serving environments can provide OpenAI-compatible APIs.

04CAPACITY

GPU Capacity

GPU capacity when you need it.

Access Australian-hosted GPU infrastructure without purchasing the underlying hardware.

  • DEDICATED

    Single-tenant GPU servers allocated to your organisation.

  • RESERVED

    Committed capacity for regular training and production workloads.

  • ON-DEMAND

    Flexible access for experimentation, bursts and development.

  • PRIVATE CLUSTER

    Isolated multi-node environments for larger AI programmes.

Commercial model

Monthly or per GPU-hour

Architectures

NVIDIA Blackwell · NVIDIA Hopper · RTX PRO

Architecture and availability are subject to deployment and capacity availability.

Register GPU Forecast
Long cold aisle of server racks with active blue status LEDs

Use cases

  • LLM training and fine-tuning
  • Inference serving
  • Computer vision
  • Research and HPC-adjacent workloads
  • Generative AI products
  • Simulation and rendering

Australian-hosted capacity planned against customer demand.

Early capacity partners receive priority consideration when new capacity is commissioned.

Start where you need us

Four services. Use one, or all of them.

Organisations engage Peregrine at the point that matters to them. You do not need to buy all four services — each stands on its own and connects when you want it to.

A

A Startup

  1. ASSESS
  2. CAPACITY
  3. OPERATE

Peregrine sizes the workload, supplies GPU capacity and operates the AI platform.

B

A University

  1. ASSESS
  2. BUILD
  3. OPERATE

Peregrine designs and commissions research GPU infrastructure the university owns, then operates the platform for its research groups.

C

An Enterprise

  1. ASSESS
  2. BUILD
  3. OPERATE
  4. CAPACITY

Peregrine designs owned infrastructure for steady workloads, operates it, and supplements peaks with Australian-hosted GPU capacity.

D

An Existing AI Company

  1. OPERATE
  2. only

The team already has GPUs. Peregrine operates the platform layer so engineers can focus on models, not clusters.

One operating model

Your infrastructure. Our infrastructure. Or both.

Peregrine's Managed AI Platform brings hardware you own, Peregrine capacity and the public cloud you already use under a single operating model — so your AI workloads run the same way wherever they land.

YOUR HARDWARE

On-premises or colocated

PEREGRINE CAPACITY

Australian-hosted GPUs

PUBLIC CLOUD

AWS · Azure · GCP

ONE OPERATING MODEL

Peregrine Managed AI Platform · Kubernetes / Slurm · scheduling · monitoring · model serving

YOUR AI WORKLOADS

Training · fine-tuning · inference · agents · research

GPU Capacity

Compute on your terms.

Four capacity models, matched to how your workloads actually run. Capacity is matched to project requirements and subject to availability.

Single GPU server node glowing softly in a rack

ON-DEMAND

Flexible GPU access for short-lived work, with capacity matched to project requirements and subject to availability.

Best for

  • Experimentation
  • Short training runs
  • Burst workloads
  • Development environments
Row of identical GPU servers stacked in one rack

RESERVED

Committed capacity over an agreed term. For sustained workloads, reserved infrastructure can provide more predictable economics than purely consumption-based compute.

Best for

  • Regular training
  • Production workloads
  • Research programmes
  • Predictable capacity
Dedicated server rack inside a locked steel mesh cage

DEDICATED

Single-tenant GPU servers allocated to your organisation, designed to support consistent utilisation and sensitive workloads.

Best for

  • Production AI
  • Sensitive workloads
  • Consistent utilisation
  • Enterprise applications
Private caged area with a cluster of interconnected racks

PRIVATE CLUSTER

Isolated multi-node GPU environments with high-speed interconnect, scoped to larger AI programmes and institutional requirements.

Best for

  • Larger AI workloads
  • Enterprise AI
  • Universities
  • Government and regulated industries

Private AI Cloud

Cloud flexibility. Dedicated infrastructure.

Deploy an AI environment designed around your organisation rather than adapting your workloads to generic infrastructure.

  • DEDICATED CAPACITY

    GPU infrastructure allocated to your organisation rather than shared with other tenants.

  • DATA CONTROL

    Data can remain within Australian-hosted infrastructure, with access and governance designed around your requirements.

  • PREDICTABLE INFRASTRUCTURE

    Known capacity, known architecture and a consistent operating model for production AI.

  • FLEXIBLE DEPLOYMENT

    Australian data-centre, private, distributed or hybrid deployment, depending on your requirements.

Design Your Private AI Cloud
Isolated secure colocation cage with two racks inside

DEPLOYMENT OPTIONS

Australian data-centre · Private · Distributed · Hybrid

Peregrine ComputeGrid

A new model for distributed AI infrastructure.

Peregrine ComputeGrid is being developed as a distributed AI infrastructure platform connecting GPU capacity across data centres, commercial sites and renewable-energy-enabled locations.

ComputeGrid is being progressively deployed as capacity is commissioned. It is designed to complement — not replace — the cloud and on-premises environments organisations already rely on.

  • Data Centres
  • Commercial Sites
  • Industrial Sites
  • Renewable Energy Sites
  • Edge Locations
Learn about ComputeGrid
  1. CUSTOMER WORKLOAD
  2. SECURE ACCESS LAYER
  3. PEREGRINE ORCHESTRATION
  4. WORKLOAD SCHEDULING
  5. GPU NODE NETWORK

Why Peregrine

Infrastructure designed around the workload — not the other way around.

Peregrine designs, specifies, sources, integrates and commissions AI infrastructure, then operates it or hands it over. The architecture follows the workload.

  1. 01

    Australian Infrastructure

    GPU capacity hosted in Australian data centres, with options for data to remain onshore.

  2. 02

    Flexible Capacity Models

    On-demand, reserved, dedicated and private cluster models, matched to how your workloads actually run.

  3. 03

    Dedicated GPU Options

    Single-tenant servers and private clusters for organisations that need isolation and consistent performance.

  4. 04

    Distributed Architecture

    ComputeGrid is being developed to connect capacity across data centres and other suitable sites.

  5. 05

    Energy-Aware Infrastructure Strategy

    Planning that considers power availability and renewable-energy-enabled locations as capacity grows.

  6. 06

    Technical Partnership

    Engineers who work with your team from planning through commissioning and ongoing operation.

Liquid coolant lines and manifold on the rear of a GPU rack
Battery energy storage racks and switchgear supporting a data centre
More about how Peregrine works

GPU Forecast

Planning GPU capacity? Tell us what you expect to need.

Peregrine uses customer forecasts to help plan future Australian GPU capacity. Registering a forecast does not create a purchase obligation. Early capacity partners receive priority consideration when new capacity is commissioned.

Compute Assessment

Let's design your compute environment.

Tell us about your workloads, capacity requirements and deployment preferences. Our team will review your requirements and discuss the appropriate infrastructure model.

We respond to all enquiries. No obligation.

FAQ

Common questions.

Short answers on GPU compute, capacity models and how Peregrine fits alongside what you already run.

View all FAQs