Skip to content

Package migrations rebuild - #12944

Open
chrabyrd wants to merge 105 commits into
dev/8.2.xfrom
cbyrd/package-migrations-rebuild
Open

chrabyrd wants to merge 105 commits into
dev/8.2.xfrom
cbyrd/package-migrations-rebuild

Conversation

@chrabyrd

@chrabyrd chrabyrd commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor

Types of changes

  • Bugfix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)

Description of Change

Issues Solved

Closes #9054

Further comments

Versions an Arches application's graphs the way Django migrations version schema. Rebuilds the approach from #12840.

An application ships graphs. When version 2 changes one, every install that already has version 1 needs its graph rows updated, its tile data brought in line, the graph republished, and its resources moved onto the new publication. This does that, with a generator, a runner, and a ledger.

The committed graph JSON under pkg/graphs/** is the desired state, the same files load_package reads. makepkgmigrations replays the app's existing package migrations into an in-memory state, projects the committed JSON into the same shape, diffs them, and writes migrations.

commands:

makepkgmigrations <app> [--name] [--dry-run] [--check]
migratepkg [<app>] [<migration>|zero] [--fake] [--plan] [--force] [--database]
showpkgmigrations [<app>] [--list] [--plan]

guards:

  • divergence. These migrations are generated against migration history, not against your database. If a curator deleted a node in the Designer, an AlterNode would update zero rows and report success. migratepkg checks the rows a plan is about to write and refuses, naming them.
  • authoring machine. If everything a plan creates is already present, it says so and tells you to record it with --fake instead.
  • faking data. Faking structural work is safe when the rows exist. Faking tile work skips the only thing that would do it, so it is refused, with the two commands that do it properly.
  • reference data. A graph points at an ontology, a template, widgets and card components. Those come from load_package. Missing ones are named before anything runs, instead of surfacing as Ontology.DoesNotExist inside Graph.publish().
  • half reversals. migratepkg <app> zero refuses up front if any migration in the plan cannot be reversed.
  • duplicate or malformed exports. Two files carrying the same graphid is an error, because the winner would otherwise depend on filename order.
  • changes that cannot be converted automatically. A datatype change, or a node moving between nodegroups, warns that stored values need a hand written RunPackagePython migration.

testing in rascolls:

Install normally, then adopt:

python manage.py setup_db --force
python manage.py packages -o load_package -s arches_rascolls/pkg -y
python manage.py makepkgmigrations arches_rascolls
python manage.py migratepkg arches_rascolls --fake
python manage.py showpkgmigrations arches_rascolls

Make a change and ship it. Add a node to an existing card in the Designer, publish, then:

GRAPH=$(psql arches_rascolls -tAc "SELECT graphid FROM graphs WHERE slug='collection_or_set';")
python manage.py packages -o export_graphs -d arches_rascolls/pkg/graphs/resource_models -g "$GRAPH"
python manage.py makepkgmigrations arches_rascolls
python manage.py migratepkg arches_rascolls --plan

Your own database already has the rows, so record the graph migration and run the data one:

python manage.py migratepkg arches_rascolls <last graph migration> --fake
python manage.py migratepkg arches_rascolls

Optional, and the path nothing exercised before. Blow the database away and let the migrations create every graph:

python manage.py setup_db --force
python manage.py load_ontology -s arches_rascolls/pkg/ontologies/linkedart
python manage.py migratepkg arches_rascolls --plan
time python manage.py migratepkg arches_rascolls

The ontology has to be loaded first. Graphs point at one by foreign key, and the pre-flight check refuses without it rather than failing inside Graph.publish().

known limitaitons:

  • card constraints are not migrated yet. ConstraintModel is not a concrete field of CardModel, so the projection does not see it.
  • migratepkg does not reindex. Reindex affected resources afterwards or search will reflect the old graph.
  • datatype changes and nodes moving between nodegroups warn rather than convert. Both need a decision a diff cannot make.
  • migrations own graphs, and nothing else. Ontologies, controlled lists, concepts, functions, plugins and widgets still come from load_package. A release that adds a node using a new controlled list has to ship the list separately, and nothing orders the two.

@chiatt chiatt left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

For string formatting - could you use f-strings or .format() for localized strings?

@chrabyrd chrabyrd changed the title Cbyrd/package migrations rebuild Package migrations rebuild Sep 18, 2026
@chrabyrd
chrabyrd added this pull request to stack #12945 September 18, 2026 20:17
@chrabyrd
chrabyrd requested a review from chiatt September 18, 2026 20:28
@CLAassistant

CLAassistant commented Sep 18, 2026 •

Copy link
Copy Markdown

CLA assistant check
All committers have signed the CLA.

@chrabyrd
chrabyrd force-pushed the cbyrd/package-migrations-rebuild branch from cea2bd4 to 6c36f26 Compare September 18, 2026 20:45
@chrabyrd
chrabyrd removed this pull request from stack #12945 September 18, 2026 20:53
@chrabyrd
chrabyrd changed the base branch from cnh/package-migrations to dev/8.2.x September 18, 2026 20:54
@chiatt
chiatt requested a review from apeters September 22, 2026 17:26
@apeters

apeters commented Sep 25, 2026

Copy link
Copy Markdown
Member

EXAMPLE COMMAND OUTPUT

CommandError: Everything these package migrations create is already here, which is what the machine they were authored on looks like. Record them instead of applying them:
  python manage.py migratepkg my_proj --fake
  my_proj.0004_graph_nodegroup_d6a16568_node_code_type_node_zip_code_and_more: NodeGroup d6a16568-204a-4adb-ba7f-12557a96da60 already exists
  my_proj.0004_graph_nodegroup_d6a16568_node_code_type_node_zip_code_and_more: Node aa2c339a-905b-43b3-bac7-26eab435ed66 already exists
  my_proj.0004_graph_nodegroup_d6a16568_node_code_type_node_zip_code_and_more: Node d6a16568-204a-4adb-ba7f-12557a96da60 already exists
  my_proj.0004_graph_nodegroup_d6a16568_node_code_type_node_zip_code_and_more: Edge 5ed88ffb-1d06-4df3-a38e-c6f2eb9b47a8 already exists
  my_proj.0004_graph_nodegroup_d6a16568_node_code_type_node_zip_code_and_more: Edge b42d847f-6865-44ca-8e39-9ab0c487152d already exists
  my_proj.0004_graph_nodegroup_d6a16568_node_code_type_node_zip_code_and_more: CardModel 40e24f6c-904a-4338-8659-4fff3ad1d44c already exists
  my_proj.0004_graph_nodegroup_d6a16568_node_code_type_node_zip_code_and_more: CardXNodeXWidget 16860ef6-42b7-4702-8191-86d54b134c8b already exists
Reconcile the graph, or re-run with --force to apply anyway.

The error messages are confusing.... For example

Everything these package migrations create is already here, which is what the machine they were authored on looks like.

Maybe better would be

It looks like these migrations have already been applied but not recorded in the system.  
Record them instead of applying them by running the following:
  
  python manage.py migratepkg my_proj --fake

Also, there are directions at the very beginning and at the end. I think all the directions should be at the end.

The ending message is also confusing given the message at the top. (Do I need to update the graph to have all updates the migration was going to apply??... isn't that the point of the migration? Or should this message even be shown if I already have the change (per the first message))

Reconcile the graph, or re-run with --force to apply anyway.

@apeters

apeters commented Sep 25, 2026 •

Copy link
Copy Markdown
Member

After the above error I ran it with the --fake flag.

python manage.py migratepkg my_proj --fake

CommandError: These package migrations change business data, which nothing else will do if they are recorded rather than run:
  my_proj.0005_data_add_node_to_tiles_aa2c339a_and_more
Fake the graph migrations only, by naming the last of them, then apply the rest:
  python manage.py migratepkg my_proj <last graph migration> --fake
  python manage.py migratepkg my_proj
Re-run with --force to record them anyway.

Again, the message is confusing (These package migrations change business data, which nothing else will do if they are recorded rather than run:)

Is the <last graph migration> the last onc successfully recorded in my system, or the last one I'm trying to run right now? Either way, I would think this message could tell me exactly what I should run.

Finally, while this command could change business data, I don't have any tiles in the system, so that warning doesn't really apply in my case.

@apeters

apeters commented Sep 25, 2026

Copy link
Copy Markdown
Member

When a migration is reverted the UI end up in this state.

Screenshot 2026-09-25 at 1 31 52 PM

@apeters

apeters commented Sep 25, 2026

Copy link
Copy Markdown
Member

It would be nice to add a message after a person calls migratepkg telling the user to reindex their data otherwise stale data could be served.

@apeters

apeters commented Sep 26, 2026

Copy link
Copy Markdown
Member

Can we get rid of the step to export graphs and just read them directly from the database?

Currently it's:

  • Update graph in the db using the designer
  • Export the graph to a file by reading from the db
  • Make a migration by reading from the file generated by the db

Maybe it could be this:

  • Update graph in the db using the designer
  • Make a migration by reading from the db (optionally updating the files on disk)

@chrabyrd

Copy link
Copy Markdown
Contributor Author

Can we get rid of the step to export graphs and just read them directly from the database?

Currently it's:

  • Update graph in the db using the designer
  • Export the graph to a file by reading from the db
  • Make a migration by reading from the file generated by the db

Maybe it could be this:

  • Update graph in the db using the designer
  • Make a migration by reading from the db (optionally updating the files on disk)

Unless I'm mistaken, we don't want that pattern. These are package migrations meant to work off the package as the source of truth. But yeah happy to chat theory around this thing 👍

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Design and implement a pattern for database migrations in a project to migrate tiles

5 participants