Release Notes

This is the list of changes to RUBEM between each release. For full details, see the commit logs on the Github page.

For a list of known issues and their fixes, visit the Github issues page.

Unreleased

The format follows Keep a Changelog.

Added

  • Added rubem calibrate: a differential evolution (SciPy) over the free calibration parameters, with w3 derived from the other two weights, minimizing 1000 (100 (1 - NSE))^2 on the station series of one output variable (arn by default) against an observed series given in the table layout of the model’s time series or as a PCRaster time series with its header. Negative, -9999 and PCRaster missing values are gaps; a spin-up window, per-run bounds (--bound), fixed parameters (--fix), the stations of the objective (--stations) and the initialization, strategy and polishing of the search are options. Every candidate runs in a fresh spawned worker through the Python API, and the run directory receives observed.csv, evaluations.csv, stations.csv, best_<variable>.csv, result.json and <config>-calibrated.json, with progress lines on the terminal. SciPy comes with the optional extra rubem[calibration]. The calibration page of the documentation describes the method, the evaluation budget, the fixed drainage network a calibration on arn needs and the process model; a dataset pytest marker with the RUBEM_DATASET_DIR convention carries a smoke test on the Ipojuca basin (#343).

  • Added rubem run --allow-blocking-problems, which runs the model even when the input validation finds blocking problems: the same checks run, every problem is still reported (the non-blocking ones as warnings, the blocking ones as errors, followed by one error line stating that the simulation continues despite them) and the run goes on instead of stopping with exit code 1. The option cannot be combined with -s, which skips the checks it reports (usage error, exit code 2), and the deprecated rubem -c <config> spelling does not take it. ModelConfiguration(..., allow_blocking_problems=False) is the library counterpart and keeps the full list in ModelConfiguration.problems; Model.from_file and Model.from_config of rubem.api take the same keyword and an isolated run rebuilds the configuration with it. rubem calibrate --allow-blocking-problems, and CalibrationSettings(allow_blocking_problems=True), do the same for a calibration, whose inputs are validated once before the search and never again by the workers; the value is recorded under settings in result.json, since the parameters it reports were fitted on inputs the validation rejected (#352).

  • Added rubem.api, the public Python surface: Model.from_file and Model.from_config load a configuration file or document, Model.run() runs the simulation in the current process and Model.run_isolated() in a fresh spawned subprocess, both returning a RunResult that lists the rasters, time series and metadata the run wrote, enumerated from the configuration. Importing the module needs neither PCRaster nor GDAL; running without them raises ImportError with the installation guidance. ConfigurationError now survives pickling with its problems (#341). The Python API page documents the stability policy, the process model and how to run in parallel.

  • Added rubem preprocess kp, which builds the Class A pan coefficient series from wind speed and relative humidity raster series and a fetch distance with the formula of the published model (supplement S29), through the shared raster I/O of the preprocessing tools; a member with a non-positive coefficient refuses the whole run, cells above the configured kp range are warned about. The model’s get_pan_coef_et_open_water_area now delegates to the same implementation (#330).

  • Added the comparison of GRID.grid (raster_info.grid_size) with the pixel size of the clone when the reference coordinate reference system is projected: the linear unit is converted to metres and a relative difference above 1e-6 on either axis blocks the run; a geographic or absent system leaves the declared value as given and logs the raster resolution (skipped with -s; #329). The user guide states that grid is the metric cell size the user asserts.

  • Added the paper conformity tests (tests/paper, marker paper): equation tests of the process functions against the journal supplement (S3, S5 to S20 and S22 to S33), model-level tests of the rules the model applies inline (open water, impervious and saturated cells, total and routed discharge, crop coefficient threshold) and an independent float64 reference of one monthly step compared cell by cell with the model on the synthetic dataset (#332).

  • Added the validation of the weighted runoff coefficient domain C_wp <= 1: for every pair of a land use class and a soil class the slope-free part of the coefficient is computed from the Manning roughness, the wilting point, the impervious and open water fractions and the weights w1, w2 and w3; a pair at or above 1 blocks the run, a pair that only the slope term can push above 1 is reported with the slope threshold, and a wilting point at or above 1 blocks (skipped with -s; #328). The user guide states the domain next to the weights.

  • Added the optional RASTERS.georeference raster whose coordinate reference system is written to the GeoTIFF outputs; the clone and the georeference must share the DEM geometry, rotated grids are refused when PCRaster maps are written, and a GeoTIFF that cannot be written is removed instead of being left half-written.

  • Added content validation of the inputs (skipped with -s): lookup tables must parse, dg, Zr, Tsat, manning and the rainy days must be positive, the rainy days must cover the twelve months, Tcc > Tw for every class; kp must be positive and NDVI below 1 in every cell, ndvi_max > ndvi_min per cell, sample identifiers contiguous from 1; the precipitation, ETP and Kp series must cover every simulated step and the NDVI and land use series the first one. Blocking problems raise ConfigurationError; kc_max < kc_min, later NDVI/land use gaps and area fractions not adding up to 1 are reported as warnings.

  • Added ModelConfigurationFile, the legacy JSON file as a validated model: the spellings found in circulating files are accepted as aliases (K_sat, T_ini, w1, kcmin, …), unknown keys are reported and ignored, duplicated keys are reported (the last value still wins), and relative paths are anchored on the directory of the JSON file (ModelConfiguration.load(path)) or on an explicit base_dir.

  • Added the model of configuration file format 1.0 (ModelConfigurationFileV1: strict keys, version, metadata, ISO simulation_period, the dated, monthly and directory raster series specifications, model_simulation_output with per-format selections) and its conversions from and to the legacy file. The format is not yet read by the loader nor exposed on the command line.

  • Resolved the raster series to one path per step through resolvers (directory, dated and monthly series; MissingStep markers instead of exceptions), used by the model for every series; the legacy directory series resolve to the same PCRaster file names as before.

  • Activated configuration format 1.0: a file with version is read as such (strict keys, duplicated keys rejected), raster series and time series are selected independently with their formats (CSV converts the .tss files, PCRasterTSS keeps them, both do both), metadata.json is written next to the outputs, rubem config schema prints the 1.0 schema by default and rubem config migrate converts a legacy file (paths rebased onto the destination, atomic write, --force to overwrite).

  • Started the preprocessing overhaul: rubem preprocess sub-command (info describes a raster), shared raster I/O with explicit contracts (atomic writes, natural ordering, geometry checks, collision detection, all-no-data policy, manifest.csv), and the legacy scripts no longer run at import time.

  • Added rubem preprocess tif2map, tif2mapseries and mapseries2tif (rubem.preprocessing.conversions): value scale by option, natural file order, PCRaster 8.3 naming, geometry checks against a clone, no-data policy, georeference for the GeoTIFF outputs; the legacy modules tif2map, tif2pcrtss and pcrtss2tif are deprecated.

  • Added rubem preprocess minmax (rubem.preprocessing.minmax_series): per-cell minimum and maximum of a raster series ignoring missing cells, with the geometry checked across the series; the legacy minmax module is deprecated.

  • Added rubem preprocess krige (rubem.preprocessing.kriging_series, optional rubem[preprocessing] extra): ordinary kriging of station series onto the clone grid, one map per step, reading the legacy matrix layout or a long step;id;x;y;value layout, with the negative-value policy (clamp by default), the kriging metric derived from the clone’s coordinate reference system and the variogram settings; the legacy kriging module is deprecated.

  • Accepted GeoTIFF input rasters and series (.tif/.tiff): rasters are read through GDAL onto the clone grid, a GeoTIFF clone sets the grid through its geometry, GeoTIFF series members are named like the model outputs, sample locations may be a GeoTIFF, and every input raster must share the clone geometry and coordinate reference system.

  • Added the spatial aggregation of the time series in configuration format 1.0 (time_series_samples.aggregation): point (as before), subcatchment (the catchment upstream of each sample over the LDD) and zones (a rasters.zones raster, ids remapped to 1..N and recorded in zones_mapping.csv); non-point tables are named tss_<variable>_<aggregation>.

  • Added the optional lai_max lookup table (TABLES.lai_max, format 1.0 lookup_tables.lai_max) with the maximum leaf area index of each land use class, selected by the constant lai_max_from_table (default false: the lai_max constant is used as before). When read, the table must be positive, at most the admissible maximum of the constant and keyed by the classes of the area fraction tables (skipped with -s); the switch without a table blocks the run even with -s, and a table without the switch is reported as ignored (#353).

Changed

  • The model overview states, next to the equations they modify, the six rules the model applies beyond the published formulation and confirmed by the model authors: the zero floor of the root zone storage, the saturated root zone of open water cells, the saturation-excess runoff SR = P - I, the cap of the open water evapotranspiration at the precipitation and the zero floor of the open water runoff, the constant impervious evapotranspiration i_imp (1 to 3 mm), and the domains kp > 0 and 0 <= C_wp <= 1 (#331).

  • The station time series the model keeps as PCRaster .tss files carry the PCRaster header (title, number of columns, timestep line and one line per station id), so PCRaster’s own tools read them; the CSV conversion reads the header, refuses a file without it or whose ids differ from the configured stations, and converts the data rows only, leaving the CSV tables unchanged (#347).

  • The impervious area interception i_imp must lie between 1 and 3 mm, the range of the published formulation (#327); the model overview and the user guide state the same range.

  • Packaged RUBEM with pyproject.toml: pip install support, the rubem console script and a single PEP 440 version source.

  • Stated the license expression consistently as GPL-3.0-or-later (the source headers’ “version 3 or any later version”) in CITATION.cff, the README and the FAQ; the license itself is unchanged.

  • Rebuilt the CI (lint, OS/Python matrix, documentation build, packaging smokes) and updated the documentation build mechanics.

  • Enabled time series per output variable and made the CSV conversion transactional over the run’s own .tss files.

  • Gave rubem.cli.main an argument list parameter, removed the PyInstaller launcher module, and made the GeoTIFF writer reject unsupported formats. Run the model with rubem or python -m rubem; executing the package directory as a script (python rubem) is no longer supported.

  • Moved the package to pathlib (enforced by ruff’s PTH and UP rules); every public path parameter accepts str or os.PathLike, and bytes paths are deprecated (accepted with a DeprecationWarning for one minor release).

  • Rebuilt the configuration value objects (SimulationPeriod, RasterGrid, CalibrationParameters, InitialSoilConditions, ModelConstants) as frozen Pydantic models with the same keywords, attributes and messages; pydantic is now a runtime dependency. New checks: the grid size must be finite with a finite square, the FPAR bounds must satisfy 0 < min < max < 1 and lai_max must be positive.

  • Rebuilt the output configuration objects as Pydantic models: OutputVariables holds OutputVariable objects (attribute access; the dictionary-style get() is deprecated), OutputDataDirectory creates the directory in ensure_exists() rather than on construction, and OutputRasterBase.from_file() reads the geometry.

  • Rebuilt the input file objects (InputRasterFiles, InputRasterSeries, InputTableFiles) as frozen Pydantic models with the same keywords, attributes and exceptions; validation problems are Problem objects and ConfigurationError carries the blocking ones.

  • Rebuilt the application settings as a plain Pydantic model (AppSettings.default() selects the PYTHON_ENVIRONMENT file at call time; the ranges singleton is gone) and made the command line report an invalid configuration with its message instead of a traceback.

  • Moved the command line to Typer: rubem run -c <config> [-s] runs a simulation and rubem config schema --format legacy prints the JSON Schema of the configuration file; rubem -c <config> still works for one minor release with a deprecation warning. typer is a runtime dependency.

  • Hardened the supply chain: every GitHub Action is pinned to a commit SHA and checked by a blocking workflow-lint job (actionlint, zizmor, pin check); checkouts no longer persist credentials; an OpenSSF Scorecard workflow publishes its results; releases are signed with Sigstore, carry a build provenance attestation, a CycloneDX SBOM of the published wheel and the conda inventory of the byte-exact environment.

  • Wrote metadata.json only after a successful format 1.0 run instead of while loading the configuration, made every GENERATE_FILE flag required again in the legacy file (as before the Pydantic rewrite), and restricted rubem preprocess krige --variogram-model to spherical, exponential and gaussian, the models the variogram fit and the interpolation share.

Fixed

  • Fixed the regression test oracle (corrected fixture inputs, structured comparators, byte-exact reproduction on a frozen environment).

  • Stopped changing the process working directory during a run; time series and raster outputs are addressed by absolute paths.

  • Exported time series only after a successful run, replaced library print() calls with log records, and fixed the output summary flags.

  • Failed clearly when the first NDVI or land-use raster cannot be read.

  • Honoured RASTER_FILE_FORMAT.map_raster_series (PCRaster maps can now be disabled; a configuration with output variables but no raster format is rejected), added the optional RASTER_FILE_FORMAT.no_data_value for the GeoTIFF series (default -9999), and matched raster series file names with the prefix taken literally.

  • Corrected the spelling of Interception.get_reflectances_simple_ratio, soil_moisture_content_wilting_point and the rubem.file._file_conversions module; the old names still work for one minor release and emit a DeprecationWarning.

  • Closed the validation gaps found in review: no-data values must fit a Float32 band, format 1.0 dates must be ISO strings, dated raster ranges are compared by month and sorted before the overlap check, lookup tables must have one key column with numeric interval bounds, ndvi_max cells equal to 1 are rejected, the blocking content rules only apply to the simulated window, the legacy file accepts the documented Kp, K_c_min and K_c_max spellings, path-like values and rejects two spellings of one key, the legacy JSON schema advertises the DD/MM/YYYY dates, nested output variables agree with tss and their field, relative series directories are frozen at construction, a time-series-only configuration is no longer reported as producing no output, and the cached application settings are read-only.

  • Reported a ValueError raised by the run or by the CSV export as an unexpected failure (with its traceback) instead of an invalid configuration, gave rubem preprocess info the same error handling as the other subcommands, and normalised bytes and bytes-valued path-like inputs in as_path() and tss2csv().

  • Preprocessing: directional maps keep fractional values, south-up or mirrored geometries are refused, geometry checks compare the coordinate reference system, mapseries2tif checks the geometry without a georeference and promotes the band type when the no-data value does not fit the source type, a stale manifest.csv or skipped member never survives a rerun, minmax refuses identical output paths, and kriging fits the variogram with the great-circle distance on geographic coordinates, rejects non-finite station cells and stations sharing a coordinate.

  • GeoTIFF inputs: .map LDD rasters are converted with pcr.ldd again, series members are checked against the clone’s coordinate reference system, members are found regardless of the extension case, flipped clone transforms are refused, sample and zone identifiers must fit a 32-bit integer, and the point-sampling field is released once the writers hold the file path.

  • CI: the Scorecard job reads the repository again, the release SBOM is generated from an environment holding only the wheel, the pre-commit file hooks run in the lint job, and .editorconfig leaves the PCRaster series members alone.

  • CI: pushes to main no longer cancel each other (one concurrency group per commit; pull request runs still supersede each other), and the Codecov statuses carry explicit thresholds (codecov.yml: the project value may move by 0.5%, the patch status is informational; the hard gate stays pytest-cov’s fail_under).

  • Outputs and inputs: a GeoTIFF output whose valid cell equals the no-data value fails the run instead of reading back as missing, a categorical GeoTIFF input (soil, land use, LDD, samples, zones) with a non-integer value or one outside the 32-bit integer range is refused at validation and at read time instead of being rounded or wrapped, a valid False cell of a boolean GeoTIFF is no longer read as missing, and the .tss to CSV conversion checks the column count of every data row.

  • Preprocessing: minmax writes the common type of the series’ bands (Float64 sources no longer overflow to inf in a Float32 output), kriging removes a stale manifest.csv before writing, and the unused --seed option of rubem preprocess krige is gone (nothing in the kriging path draws random numbers, and it reseeded the process’s global NumPy generator).

  • Configuration: the logging settings of the cached AppSettings are read-only all the way down (get_setting and model_dump hand out plain copies), and the per-variable output flags follow Pydantic’s bool parsing ("false" disables a variable instead of enabling it; an unparsable string is rejected).

  • Interception: positive monthly precipitation reaches the denominator of the interception-rate equation unchanged; the 1e-5 guard against a zero precipitation was being added to every value (#319). The golden fixtures were regenerated (see tests/fixtures/AUDIT.md).

  • Evapotranspiration: a cell whose NDVI equals 1.1 * NDVI_min takes the kc_min branch of the crop coefficient, as documented, instead of a crop coefficient of zero (#320). The reference dataset has no cell on the threshold, so the golden fixtures are unchanged.

  • Baseflow: the recession baseflow is limited to the water available in the saturated zone (TU_S of the previous step plus the recharge), so the saturated-zone storage can no longer become negative (#322).

Removed

  • Replaced the PyInstaller bundles with sdist/wheel distributions verified by the release pipeline.

Version 0.9.0-beta.3

Date: Mar 21, 2024

Version 0.2.3-beta.2

Date: Jan 24, 2024

Version 0.2.2-beta.1

Date: May 17, 2023

Version 0.1.3-alpha

Date: March 23, 2022

Version 0.1.0-alpha

Date: November 23, 2021