Skip to content

Unable to run MPI within Openshell #3019

Description

@RemiLehe

User Story

As a computational physicist, I want the agent to be able to run MPI-enabled applications within the Openshell sandbox, so that the agent can help me test/develop/debug MPI-enabled scientific software for High-Performance Computing (HPC).

Problem Statement

When running an MPI-enabled application within the sandbox, the application fails at initialization (MPI_Init) with the following message:

Abort(404899215): Fatal error in internal_Init_thread: Other MPI error, error stack:
internal_Init_thread(71).: MPI_Init_thread(argc=(nil), argv=(nil), required=3, provided=0x7ffda19f22f8) failed
MPII_Init_thread(261)....: 
MPID_Init(549)...........: 
MPIDI_UCX_init_local(234): 
MPIDI_UCX_init_worker(86):  ucx function returned with failed status(ucx_init.c 86 MPIDI_UCX_init_worker Shared memory error)
Abort(539116943): Fatal error in internal_Init_thread: Other MPI error, error stack:
internal_Init_thread(71).: MPI_Init_thread(argc=(nil), argv=(nil), required=3, provided=0x7ffc5cc9ebc8) failed
MPII_Init_thread(261)....: 
MPID_Init(549)...........: 
MPIDI_UCX_init_local(234): 
MPIDI_UCX_init_worker(86):  ucx function returned with failed status(ucx_init.c 86 MPIDI_UCX_init_worker Shared memory error)

Impact / Why This Matters

Having the ability for the agent to run MPI-enabled application is essential for it to help with software development for HPC. Without it, the agent can modify the code to implement new features, but is then unable to test the changes properly.

The current work-around is to do agent-assisted software development outside of openshell, i.e. without the safety guardrails that Openshell offers.

Acceptance Criteria

The bug would be considered fixed if I can run an MPI-enabled application to completion (without errors) from within an openshell sandbox (see reproducer below).

Reproduction Steps

This shows the error with a minimal MPI application: in this case a Python script using mpi4py (but the bug is not specific to Python or mpi4py, and I also observed it in C++ applications using MPI)

  1. Create a file Containerfile with the following content, so that mpi4py is available in the image (installed via conda):
FROM ghcr.io/nvidia/openshell-community/sandboxes/base:latest

RUN curl -L -O https://github.com/conda-forge/miniforge/releases/latest/download/Miniforge3-Linux-x86_64.sh \
    && bash Miniforge3-Linux-x86_64.sh -b -u -p /sandbox/conda \
    && rm Miniforge3-Linux-x86_64.sh \
    && /sandbox/conda/bin/conda clean --all -y \
    && /sandbox/conda/bin/conda init bash

ENV PATH=/sandbox/conda/bin:$PATH

RUN conda install -y -c conda-forge mpi4py \
    && conda clean --all -y
  1. Build the image locally: podman build -t localhost/test:latest .

  2. Create a sandbox using this image: openshell sandbox create --from localhost/test

  3. From inside the sandbox, run mpirun -np 2 python -c "from mpi4py.MPI import COMM_WORLD"

Note that the same error does not occur when running this simply through podman instead of Openshell. More specifically:

podman run -it localhost/test:latest

and then from within the image:

mpirun -np 2 python -c "from mpi4py.MPI import COMM_WORLD"

works as intended, showing that the issue is specific to Openshell, not to the podman image.

Environment

  • openshell 0.0.80
  • Ubuntu 26.04
  • Runtime: podman version 5.7.0

Metadata

Metadata

Assignees

No one assigned

    Labels

    state:triage-neededOpened without agent diagnostics and needs triage

    Type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions