Skip to content

REST please export PoolManager’s pool status #8157

Description

@calestyo

Hey.

(for me, having #8126 would still be the most helpful right now)

This splits out another wish from #8009 into it's own issue, simply to keep things more digestible.

In admin, PoolManager has e.g.:

> \sp psu ls pool -l
lcg-lrz-dc01_0  (enabled=true;active=0;rdOnly=false;links=0;pgroups=1;hsm=[];mode=enabled)
 linkList   :
lcg-lrz-dc01_1  (enabled=true;active=24;rdOnly=false;links=0;pgroups=1;hsm=[];mode=enabled)
 linkList   :
lcg-lrz-dc01_10  (enabled=true;active=0;rdOnly=false;links=0;pgroups=0;hsm=[];mode=enabled)
 linkList   :
lcg-lrz-dc01_11  (enabled=true;active=0;rdOnly=false;links=0;pgroups=0;hsm=[];mode=enabled)
 linkList   :

Now mode is already exported by REST as part of the /pool/{name}/usage resource, but - as far as I can see - the others aren’t.

In particular, it wish for enabled, active and rdOnly to be exported (please read below how I think this should be done).

Why?

First I'd say at least active is a good indicator of whether a pool is still up. I think the default heartbeat frequency is 30s, so everything way above means something is fishy.

Also, there's my comment. The problem with e.g. /pools/ and /cells is, that, if the pool "stops" (or some cell), it's removed from that resource (like completely).
I'm not exactly sure what "stop" here means. E.g. if I do a mere pool disable -strict, the pool cell is still running and the pool is still in /pools, but if I systemctl stop it, like in the comment above, it's gone (even if it didn't properly sign off and e.g. PoolManager still shows it as mode=enabled (again, see the comment).

After some thinking I would however say that this is actually the right behaviour: if a pool or cell is gone, it shouldn't be in /pools/ respectively /cells anymore.

However, one still wants some indication that a pool, which perhaps should be there (and just isn't due to failure), is no longer.

And conceptually I think PoolManager is the right place, because it knows all the pools that have been created psu create pool, and that entry in the psu doesn't go away, just because I stop the pool.

Long story short, it seems to me, if one wants to know whether a pool that should be there is no longer, PoolManger is the right entity to ask, which is why I'd like that to be exported in the REST. :-)

Now, how should this (IMO) be exported:

  • First, we may have many PoolManagers, and the values may differ between them (like activeness)... so we need the information for every PoolManager
  • What's IMO best is: having e.g. /poolmanager/{name|*}/status, if there's a name, return only the status for that PoolManager, if it's the literal * return all, but please don't use a JSON array, but a JSON object with the full address of the PoolManager as key (far easier to work with).

Looking at:

public String toString() {
return getName() + " (enabled=" + _enabled + ";active=" + (_active > 0 ? (getActive()
/ 1000) : "no") + ";rdOnly=" + isReadOnly() + ";links=" + _linkList.size()
+ ";pgroups=" + _pGroupList.size() + ";hsm=" + _hsmInstances.toString() + ";mode="
+ _mode + ")";
}

  • enabled and rdOnly, can only be true or false ... JSON has booleans, so better these than strings
  • active is weird... it can also be no (see psu set active <pool>|* [-no]) but I don't understand what that means
    IMO, the JSON should export:
    • one field active, which is again a boolean
    • And the information of the last_heartbeat_timestamp... but as UNIX epoch timestamp (which dCache anyway uses internally)... not as "seconds since". The latter has many problems... i.e. when caching is done or when getting the data takes a while... it still says e.g. 10s but that's of course no longer true. Also, if dCache internally has he information with higher details than seconds (there's a / 1000 in the code)... then please export as detailed as possible.
  • I don't know how expensive gathering this data is... but having it cached would of course mean that it's often outdated... in particular for the heartbeat, which is normally <30s... the caching would obviously make it far less useful.

Should mode= be included? Well, in principle I'd say no, because it's not really a canonical value of the PoolManager... but since the pool disappears from /pools once it's gone... I don't now what's best.
I mean even with all this data exported (and even with the current data from psu ls pool -l) there's still the problem: when should one actually consider a pool dead?
I mean in particular if - as shown in #8156 - the mode= may not be updated anyway, the only thing that remains is a too high active= value.

Thanks,
Chris

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions