Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
173 changes: 173 additions & 0 deletions docs/kb/reference/referential-integrity.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,173 @@
---
title: Referential Integrity
description: What FOG 1.6 does when you delete something other records point at
context_id: referential-integrity
aliases:
- Referential Integrity
- Foreign Keys
tags:
- 1_6-changes
- database
- storage
- images
- hosts
- management
---

# Referential Integrity

Up to and including FOG 1.5, nothing in the database itself recorded that one
record pointed at another. Cleaning up after a delete was the job of the PHP
that ran the delete, which meant every delete path had to remember every
dependent table — and a path that forgot left rows behind pointing at
something that no longer existed. Nobody found those rows until one of them
did something visible: a host in a group that had been deleted, a task
against an image that was gone, a storage node in no group at all.

FOG 1.6 declares those relationships in the database. The rules below are
enforced by MariaDB on every delete, whatever performed it — the web UI, the
API, a plugin, or a query someone typed by hand.

>[!note]
>Nothing here changes what you can *create* or *edit*. It only changes what
>happens to related records when something is deleted, and which deletes are
>refused outright.

## The three outcomes

Every relationship resolves to one of three behaviors:

| Outcome | What happens | Used for |
|---|---|---|
| **Cascade** | The dependent rows go too, in the same operation | Things owned by the record — a host's MAC addresses, a group's membership rows |
| **Clear the reference** | The dependent row survives with the reference emptied | Things that *mention* the record — a host's assigned image |
| **Refuse** | The delete is rejected and nothing changes | Configuration something else is actively relying on |

**The audit trail and history take none of these.** Audit and history rows
deliberately outlive their subject: the record of who deleted a host is worth
nothing if it disappears when the host does. Those rows stay, and they keep
naming the id that used to exist.

## What each delete now does

### Deleting a host

Everything belonging to that host goes with it, in one step: its MAC
addresses, group memberships, snapin and printer assignments, inventory,
module status, screen and auto-logout settings, power management, site
membership, its tasks and snapin jobs, and its location and OU associations
if those plugins are installed.

Its **history and audit rows stay.**

### Deleting an image

- **Hosts assigned that image are unassigned** — the host survives with no
image, rather than pointing at one that is gone.
- Scheduled tasks and any running task lose the image reference the same way.
- Storage-group associations for the image are removed, as are Windows key
associations if that plugin is installed.

### Deleting a storage group

**This is refused** while any of the following still points at the group:

- **storage nodes** that belong to it
- **pending file deletions** queued against it
- a **location** naming it, if that plugin is installed

Move or delete those first, then delete the group. Snapin and image
associations to the group are *not* a blocker — those are removed with it,
and a running multicast session on the group ends with it.

>[!warning]
>Before 1.6 this delete was allowed and silently orphaned whatever was
>pointing at the group. On one real installation, deleting a single storage
>group left three storage nodes belonging to nothing.

### Deleting a storage node

Not refused. Anything referring to the node loses the reference and carries
on: a running multicast session drops its sender node, tasks drop their node
references, and a **location** that named that specific node falls back to
choosing the best node in its storage group — which is the same thing a
location does when no specific node is set.

### Removing a storage node from its group

**A storage node must belong to a storage group.** A group with no nodes is
fine; a node in no group is not — it is invisible to replication and to
multicast, and nothing will ever assign it work.

There is therefore no "remove from group" operation any more. **To move a
node, assign it to the group you want it in** — that moves it in one step,
and leaving its old group empty is not a problem.

### Deleting an image type, an OS, or a task state

**Refused** while any image or task is using it. These are the fixed lists
FOG's own behavior keys off; deleting one out from under a record in use
would leave that record undescribable.

### Deleting a user or a role

Everything granted to or through them goes: site memberships, role
assignments, group memberships, API tokens, stored credentials, and any
LDAP or OIDC grants and identities.

An OIDC **identity** — the record that a particular external account *is* a
particular FOG user — goes with either end. That is deliberate: left behind,
the next user created could inherit someone else's sign-in binding.

## What a refused delete looks like

The refusal comes from the database, and FOG surfaces the database's own
message. It names the constraint, and the constraint is named after the table
and column that is still pointing at the record:

```
Cannot delete or update a parent row: a foreign key constraint fails
(`fog`.`nfsGroupMembers`, CONSTRAINT `fk_nfsGroupMembers_ngmGroupID`
FOREIGN KEY (`ngmGroupID`) REFERENCES `nfsGroups` (`ngID`))
```

Read it as: **`nfsGroupMembers` still has rows in this group** — a storage
node has not been moved out. `fk_<table>_<column>` always identifies what is
holding the record.

Nothing was changed when a delete is refused. Clear the thing named in the
message and try again.

## What happens on upgrade

The upgrade to 1.6 has to make the existing data consistent before it can
declare any of these rules, because the database will not accept a rule that
the rows already there break.

- **Rows pointing at something that no longer exists are cleaned up once.**
Where the row can survive with an empty reference it is emptied; where it
cannot, it is removed. On a healthy installation this finds nothing.
- **What it did is recorded in the audit log**, with a per-table count, under
the type `schema.orphan.sweep`. A silent cleanup would be worse than no
cleanup — if rows were removed, you can see how many and from where.
- **Columns that spelled "no reference" as `0` now use an empty reference
instead.** Nothing about how FOG reads them changed; `0` and empty behave
identically everywhere they are read.

>[!note]
>If a rule cannot be applied because of data the upgrade could not safely
>decide about, the upgrade **does not fail.** The rule is skipped, the reason
>is written to the web server's error log, and FOG carries on behaving as it
>did in 1.5 for that one relationship. Fixing the data and running the
>upgrade check again applies it.
>
>The one that occurs in practice is a **storage node sitting in no group**,
>from before this was prevented. Nothing can guess which group it belongs in,
>so assign it to one and the rule applies on the next upgrade check.

## Where the detail lives

The full classification of every relationship, the reasoning behind each
choice, and the measurements behind them are in the `fogproject` repository:
`docs/development/foreign-keys.md` and
`docs/adr/0031-referential-integrity-is-declared-in-the-database.md`.
14 changes: 14 additions & 0 deletions docs/management/web/images.md
Original file line number Diff line number Diff line change
Expand Up @@ -190,3 +190,17 @@ a great option to save image space and network transfer volume.
create a folder structure like so: /images/SampleXPImage
- Drop your image file into the folder (be sure it's named
the same as image name above)

## Deleting an Image object (1.6)

Deleting an image definition no longer leaves anything pointing at it:

- **Hosts assigned that image are unassigned.** The host survives with no
image rather than naming one that is gone.
- Scheduled tasks and any running task lose the image reference the same way.
- The image's storage-group associations are removed with it, as are Windows
key associations if that plugin is installed.

>[!note]
>This removes the image *definition* from the FOG database. It does not
>delete the image files on the storage node — see [[kb/reference/referential-integrity|Referential Integrity]] for the full set of delete rules.
20 changes: 20 additions & 0 deletions docs/management/web/storage-node.md
Original file line number Diff line number Diff line change
Expand Up @@ -78,6 +78,26 @@ tags:
>[!note]
>Master node status also decides which node transmits a multicast session, so a site whose node is not a master cannot serve multicast locally even with the image replicated to it. If you multicast to more than one site, see [[management/web/multicast#Multicast across multiple sites|Multicast across multiple sites]].

## Moving and deleting nodes (1.6)

A storage node **must** belong to a storage group. A group with no nodes is
fine; a node in no group is not — it is invisible to replication and to
multicast, and nothing will ever assign it work.

- **To move a node, assign it to the group you want it in.** That moves it in
one step. There is no separate "remove from group" operation, because a
node with no group is not a state FOG has.
- **Deleting a storage group is refused while nodes still belong to it**, and
likewise while a location names it or file deletions are queued against it.
Move the nodes out first. Before 1.6 this delete was allowed and silently
left those nodes belonging to nothing.
- **Deleting a storage node is not refused.** Anything referring to it loses
the reference and carries on — a running multicast session drops its sender
node, and a location that named that specific node falls back to choosing
the best node in its group.

For the full set of rules and what a refused delete looks like, see [[kb/reference/referential-integrity|Referential Integrity]].

## Including multiple PXE / TFTP servers

- A traditional Master Storage Node, [as described above](https://wiki.fogproject.org/wiki/index.php?title=Managing_FOG#Adding_a_Storage_Node) only provides File Storage redundancy. While this can help increase multicast throughput on a single network, all the machines under FOG management must be within the same subnet/VLAN so that DHCP broadcast requests can be directed to the Main server. (see note below)
Expand Down