Changelog
peskas.zanzibar.data.pipeline 4.9.0
Bug Fixes
FishBase releases are now pinned, and a missing coefficient fails the run:
rfishbasereads a remote parquet dataset over the network, so an unpinned"latest"let a new FishBase release reach the pipeline the moment a container was rebuilt, with no code change. Release 26.06 dissolvedCaesionidaeintoLutjanidaeandScaridaeintoLabridae— both family names survive with zero species in them — so any taxon named after one expanded to nothing, got no length-weight coefficients, and weighedNA, which sums to zero. The releases are now pinned per server ininst/config.ymlundermetadata:fishbaseand threaded through all fiverfishbasereads ingetLWCoeffs(), which previously could mix snapshots within a single run.rfishbaseis additionally pinned to 5.0.1 in both Dockerfiles as the last install step, becauseremotes::install_local(dependencies = TRUE)upgrades it otherwise.assert_taxa_coverage()fails the run when a taxon resolves to no coefficients: previously a taxon that matched nothing was dropped in silence and the pipeline stayed green while publishing a hole.Unmatched taxa are now logged:
match_species_from_taxa()dropped any name that matched no species without a warning.SeaLifeBase routing corrected:
process_species_list()routed only ISSCAAP groups 57, 45, 43, 42 and 56 to SeaLifeBase, sending sea cucumbers, gastropods, oysters, mussels, scallops and mantis shrimp to FishBase, where they matched nothing and were dropped. Routing is now ISSCAAP >= 40.Species names ending in “idae” are no longer read as families: the rank test placed the family suffix before the species test, so
Haliotis midaeandJordanella floridaewere searched as families and matched nothing.-
Length-type conversion recovers 16 taxa that weighed
NA(get_length_conversions(),convert_lw_to_tl()): FishBase tags every published length-weight pair with the length type the original study measured, and for tunas, billfish and several carangids that is fork length.get_length_weight_batch()kept onlyType == "TL", so those taxa got no coefficients at all and every length-measured catch row of them weighedNA. The conversions are published data in FishBase’s POPLL table, which this pipeline never read;coasts::get_taxa_morphometrics()already does. Reading it restatesaon a total-length basis (a_TL = a * ratio^b,bunchanged) and recoversALB BET BLM BUM CJC EWM FLY LJK LTQ LWO MLS NAB NXM NXP QJR SWO. Coverage goes from 98 to 114 of 134 codes, and none of the 98 coefficients that already resolved changed — the conversion is applied only to taxa that would otherwise have nothing.Validated against the 632 species carrying both a native TL pair and an FL one: converting halves the median error in predicted weight (15.6% against 27.5% for using the FL pair as-is), and against the 1,186 with both TL and SL pairs it cuts it fourfold (16.8% against 70.6%). Ratios are physically sensible (median FL/TL 0.962, SL/TL 0.831). Compared against FishBase’s independent TL-basis Bayesian estimates for the recovered taxa, converting is closer than raw on 10 of 15.
Retired survey codes are remapped on read (
reshape_catch_data()): editing a KoBoToolbox form’s choice list does not rewrite submissions already collected, so a retired code persists in historical data indefinitely.AHI(171 rows) andBFL(6 rows) had additionally been removed from the Airtable taxa table, which orphaned them —map_surveys()left those rows withNAscientific name and alpha3 code. They now remap toBAF(Ablennes hians) andTEI(Pterocaesio pisang), joining theTUNandSKHremaps that were already there.The
MACcorrection now runs early enough to matter:MACwas wrongly offered under the sharks-and-rays group in the form, and all 150 rows carrying it arefish_group == "SR"with a median length of 85 cm — eagle rays, not the Atlantic mackerel ASFIS mapsMACto. TheSR/MAC→AQXrule existed inpreprocess_wf_surveys()but ran aftercalculate_catch(), so the rows were weighed asMAC(which resolves to no coefficients, givingNA) and only relabelled afterwards. Moving it intoreshape_catch_data()lets them pick up theAQXcoefficients: 137 of 190AQXrows now carry a weight, median 34.9 kg, where the 150 previously carried none.Morphology bounds no longer silently disable length validation:
min(CommonLength, na.rm = TRUE)returnsInffor a taxon whose matched species all lack that field, and the permissiveness step then computedInf - 0.75 * Inf=NaN. Every comparison againstNaNisNA, whichcase_when()treats as no-match, so alert codes 3 and 4 never fired for those taxa — a missing bound was indistinguishable from a passed check.safe_min()now yieldsNArather thanInf, and because FishBase populatesCommonLengthfor only 10% of species against 91% forLength, missing values are estimated as0.625 * Length— the median ratio across the 3,748 species carrying both (common_length_ratio()). All 128 taxa with morphology now have usable bounds, against 20 previously broken. Expect a wave of new length alerts on the first run: those records were never checked before.-
Search-name aliases fix six taxa the ASFIS names could not match (
taxa_search_aliases(),apply_taxa_aliases()): a handful of ASFIS reference names match nothing in the taxonomic backbone, so the taxon is dropped and every catch row of it weighsNA.CLPis namedClupeidae, which FishBase emptied of Indo-Pacific species in 2022 — it now searchesDorosomatidae, matching the row already in Timor’s equivalent table. The rest are synonyms that have moved on:ESRto Stolephorus commersonnii,RPOto Parupeneus macronemus,LZVto Ellochelon vaigiensis,OQCto Octopus cyanea, andVMX— Valamugil, a genus the backbone no longer carries — to Osteomugil and Moolgarda. The table carries an explicitrankbecause the suffix rules inprocess_species_list()cannot recognise a bare genus name.CRA(“marine crabs nei”, Brachyura) is deliberately not aliased: it is an infraorder, and SeaLifeBase carries no rank between order Decapoda and family, so choosing a target means deciding which crab families Zanzibar lands. OQCnow gets the octopus mantle-length conversion:OCZ(Octopus spp) was special-cased in three places — an ML-only coefficient filter, the arm-span-to-mantle/5.5conversion incalculate_catch(), and themin_lengthfloor — butOQC(Octopus cyaneus) was not, despite being the larger of the two in the data at 694 rows with 484 lengths, median 85 cm.OQCresolves to a mantle-length pair, so applying it to arm-span unconverted weighed a single octopus at 263 kg instead of 2.43 kg. This was latent whileOQChad no coefficients and would have gone live with the alias above.Fixed a duplicate join key for
FLY:preprocess_wf_surveys()appended a hardcoded flying-fish coefficient unconditionally. Now that the conversion recovers Exocoetidae pairs,getLWCoeffs()returns aFLYrow of its own, and two rows on the same key would have doubled every flying fish catch record incalculate_catch(). The manual value now replaces rather than appends, and stays authoritative: FishBase’s recovered pairs weigh a 30 cm flying fish at 498 g against the hardcoded 202 g, and changing that is a separate decision.Removed a dead fallback in
preprocess_wf_surveys(): thetryCatcharoundgetLWCoeffs()readinst/length_weight_params.rds, which is not in the package, so the fallback could only ever fail — while hiding the original error behind it.
Known Issues
Measured 2026-09-06 against FishBase 25.04 / SeaLifeBase 24.07 over the live KoBo data: 122 of 134 codes resolve length-weight coefficients, up from 98 before this release. The other 12 form the documented baseline in assert_taxa_coverage(), so any new loss fails the run. CJX and PWT are deliberately not in it — they resolve at 25.04 and are the two codes that break at 26.06, so a release move fails the check.
-
Wrong reference name (2) —
MAE,TAGname species absent from FAO 51. -
No published coefficients (1) —
GQT(Plectorhinchus gaterinus) occurs in FAO 51 but FishBase carries no length-weight pair for it in any length type. Nothing to convert, nothing to alias. -
No convertible length type (5) —
KAK,LHV,RMB,RTY,SSP. -
Not a taxon, or a rank the backbone omits (4) —
MZZ,UNKN,UNK, andCRA(the infraorder Brachyura).
peskas.zanzibar.data.pipeline 4.8.0
New Features
-
Gleaning Survey Data Integration: Added support for KoBoToolbox gleaning survey data collection and processing
- New survey type (
gleaning) is now ingested from KoBoToolbox, preprocessed, and validated alongside existing catch surveys - Gleaning data captures informal fisheries activity not covered by formal catch surveys, providing a more complete picture of small-scale fishing effort
- Gleaning submissions are now included in the unified validation pipeline with tailored quality checks
- New survey type (
peskas.zanzibar.data.pipeline 4.7.0
New Features
- Unified API export across survey programs: The raw and validated API exports now combine data from both WCS and WorldFish surveys into a single output file. Previously only WorldFish data was exported; WCS trips are now included alongside them with a consistent set of fields.
Improvements
-
WCS validation enhancements:
- Added two new quality checks: one that catches contradictions between bucket count and bucket weight (e.g. buckets recorded but no weight, or vice versa), and one that flags implausible negative values in catch measurements
- Price and revenue validation thresholds corrected to Tanzanian Shilling values (previous values were in Mozambican metical)
- Submissions recorded before 2020 are now excluded from the validated dataset
- Catch outcome is now correctly carried through to the validation step
- WCS price calculation corrected: Catch prices are now split proportionally across species within a trip. Missing species prices now fall back to the species-level median rather than being left empty.
peskas.zanzibar.data.pipeline 4.5.0
Major Changes
-
Adopted
coastsas the shared multicountry analytics engine: Aggregated data summarization and dashboard export are now delegated toWorldFishCenter/peskas.coasts(dev branch). This centralizes the logic for producing monthly, taxa, district, and gear summaries — as well as fishery metrics — across all Peskas country deployments (Zanzibar, Kenya, Mozambique), ensuring consistent outputs and a single place to maintain and improve the shared pipeline logic.- Added
coaststoImportsandRemotesinDESCRIPTION - Added
remotes::install_github("WorldFishCenter/peskas.coasts", ref = "dev")to bothDockerfileandDockerfile.prodso the image ships the package - Pipeline steps that previously used local
summarize_data()andgenerate_fleet_analysis()now call the equivalentcoasts::functions, passingpackage = "peskas.zanzibar.data.pipeline"so they read the country-specificinst/conf.yml
- Added
peskas.zanzibar.data.pipeline 4.4.0
Improvements
-
Standardized configuration structure: Replaced
inst/conf.ymlwith a unified multi-country template harmonized across all Peskas deployments (Zanzibar, Kenya, Mozambique). Key structural changes:- Survey credentials moved from
surveys.*into a new top-levelingestion.*section - Stage keys shortened (
raw_surveys→raw,preprocessed_surveys→preprocessed, etc.) - Source names shortened (
wcs_surveys→wcs,wf_surveys_v1→wf_v1, etc.) - MongoDB structure reorganized: connection strings under
connection_strings.*, databases underdatabases.*, collections key pluralized,portalrenamed todashboard - Airtable config moved from top-level
airtable.*tometadata.airtable.* - All R code updated to use the new config paths
- Survey credentials moved from
-
summarize_data()correctness and clarity fixes:- Fixed data quality bug:
taxa_summarieswas incorrectly summing trip-total catch kg per taxon instead of the actual per-taxon catch weight — values are now correct - Fixed
districts_summariesround-trip pivot: data is now stored wide and pivoted once inexport_wf_data()instead of pivot-long → store → pivot-wide → pivot-long - Fixed
gear_summariescomplete()running afterpivot_longerwith wrong column names in the fill list; now completes before pivoting across all gear × district × month combinations - Replaced fragile
across(everything(), first)pattern in trip-level collapse with explicitslice(1), which is clearer and avoids silent column overwrites - Fixed inconsistent
na.rmusage in gear summaries
- Fixed data quality bug:
-
Removed dead code: Deleted
create_geos()andcreate_geos_v1()fromexport.R; the geographic summary logic had already been inlined intoexport_wf_data()
peskas.zanzibar.data.pipeline 4.3.0
New Features
-
Survey-GPS Trip Matching Pipeline: Added comprehensive fuzzy matching system to link catch survey records with GPS trip data
- New
merge_trips()workflow function for Kenya and Zanzibar sites -
match_surveys_to_gps_trips(): Universal two-step matching (surveys → registry → trips) - Fuzzy string matching using Levenshtein distance on registration numbers, boat names, and fisher names
- Conservative one-trip-per-day constraint to avoid ambiguous matches
- Support for both explicit device registries (Kenya) and implicit registries built from trip data (Zanzibar)
- Configurable matching thresholds (registration: 15%, names: 25% difference allowed)
- Exports merged dataset with matched pairs plus all unmatched surveys and trips
- Helper functions:
standardize_column_names(),clean_matching_fields(),build_registry_from_trips()
- New
Improvements
-
PDS Data Ingestion:
- Updated
ingest_pds_trips()to load device registry from cloud storage instead of Airtable metadata - Added device info retrieval and proper filtering for Zanzibar devices
- Improved configuration variable naming (pars → conf)
- Updated
-
GitHub Actions Workflow:
- Added new
merge-tripsjob to automated pipeline - Runs after survey preprocessing and before summarization
- Integrated with production environment configuration
- Added new
peskas.zanzibar.data.pipeline 4.2.0
New Features
-
API Data Export Pipeline: Added new
export_api_raw()function to export raw preprocessed survey data in API-friendly format- Exports raw/preprocessed trip data (before validation) to cloud storage
- Part of a two-stage API export pipeline (raw and validated exports)
- Transforms nested survey data into flat structure with standardized trip-level records
- Generates unique trip IDs using xxhash64 algorithm
- Integrates with Airtable metadata for form-specific asset lookups
- Exports versioned parquet files to
zanzibar/raw/path for external API consumption - Includes comprehensive output schema with 14 standardized fields (trip_id, landing_date, gear, catch metrics, etc.)
-
Airtable Integration: New helper functions for managing Airtable metadata and form configurations
-
get_airtable_form_id(): Retrieves Airtable record IDs from KoBoToolbox asset IDs -
airtable_to_df(): Downloads complete Airtable tables with automatic pagination handling -
get_writable_fields(): Identifies updatable fields in Airtable tables (excludes computed fields) -
update_airtable_record(): Updates individual records with field validation -
bulk_update_airtable(): Batch updates multiple records efficiently (up to 10 records per request) -
device_sync(): Synchronizes GPS device metadata between Airtable and MongoDB
-
Improvements
-
Configuration Enhancements:
- Added
apiconfiguration section for trip data exports with separate raw/validated paths - Configured cloud storage paths for API exports (zanzibar/raw, zanzibar/validated)
- Added Airtable base ID and token configuration for metadata management
- Enhanced
options_apistorage configuration for peskas-coasts bucket
- Added
-
GitHub Actions Workflow:
- Added new
export-api-datajob to automated pipeline workflow - Integrated Airtable authentication with GitHub Secrets (AIRTABLE_TOKEN, AIRTABLE_BASE_ID_FRAME, AIRTABLE_BASE_ID_ASSETS)
- Configured API export job to run after survey preprocessing step
- Added production environment configuration for API data exports
- Added new
-
Code Quality:
- Improved documentation with comprehensive roxygen2 comments for all new functions
- Added detailed examples and cross-references in function documentation
- Enhanced error handling and input validation in Airtable operations
- Implemented proper cleanup of temporary local files after cloud uploads
peskas.zanzibar.data.pipeline 4.1.1
Major Changes
-
Streamlined Validation Workflow: Replaced KoboToolbox API updates with direct MongoDB storage to improve performance.
- New
export_validation_flags()function exports validation flags directly to MongoDB - Validation status queries now only identify manually edited submissions, not update them
- Disabled
sync_validation_submissions()workflow steps in GitHub Actions - Significantly reduced pipeline execution time by avoiding slow KoboToolbox API calls
- New
Improvements
-
Validation System:
- Validation functions now preserve manual human approvals while updating system-generated statuses
-
Code Quality:
- Fixed SeaLifeBase API calls by pinning to version 24.07 to avoid server errors
- Standardized function parameter formatting across validation and preprocessing modules
peskas.zanzibar.data.pipeline 4.1.0
Major Changes
-
Integration of New KoBoToolbox Survey Form Version: Added support for a new version of the WorldFish survey form (
wf_surveys_v2) alongside the existing form (wf_surveys_v1). Data from both survey versions is now processed together in the preprocessing pipeline and handled properly throughout the validation workflow.
Improvements
-
Multi-Asset Validation Support:
- Updated validation system to query approval statuses from both survey form versions
- Enhanced
validate_wf_surveys()andsync_validation_submissions()to handle submissions from multiple KoBoToolbox assets - Ensured manually approved submissions from either form version are protected from automated flagging
-
Configuration Updates:
- Added configuration for the new survey form version with shared credentials
- Cleaned up redundant configuration entries
- Updated code references to use versioned asset configurations
peskas.zanzibar.data.pipeline 4.0.0
Major Changes
- Fleet Activity Analysis Pipeline: Introduced a comprehensive pipeline for estimating and analyzing fishing fleet activity using GPS-tracked boats and boat registry data. This includes new functions for preparing boat registries, processing trip data, calculating monthly trip statistics, estimating fleet-wide activity, and calculating district-level total catch and revenue.
-
New Modeling and Summarization Functions:
-
prepare_boat_registry(): Summarizes boat registry data by district. -
process_trip_data(): Processes trip data with district information and filters outliers. -
calculate_monthly_trip_stats(): Computes monthly fishing activity statistics by district. -
estimate_fleet_activity(): Scales up sample-based trip statistics to fleet-wide estimates. -
calculate_district_totals(): Combines fleet activity and catch data for district-level totals. -
generate_fleet_analysis(): Orchestrates the full analysis pipeline and uploads results. -
summarize_data(): Generates and uploads summary datasets (monthly, taxa, district, gear, grid) for WorldFish survey data.
-
-
Enhanced Data Export and Integration:
-
export_wf_data(): Exports summarized WorldFish survey data and modeled estimates to MongoDB, including new geographic regional summaries. -
create_geos(): Generates geospatial regional summaries and exports as GeoJSON for spatial visualization.
-
- Expanded Documentation: New and updated Rd files for all major new functions, with improved examples and cross-references.
Improvements
- Consistent Time Series and Grouping: All summary tables (taxa, districts, gear) now include a ‘date’ (monthly) column and are grouped by month, with missing months filled as NA for consistent time series exports.
-
Parallel Processing: Improved use of parallelization (via
futureandfurrr) for validation and summarization steps, enhancing performance for large datasets. -
Data Quality and Validation:
- Enhanced filtering and validation of survey data before summarization and export.
- Improved handling of flagged/invalid submissions.
peskas.zanzibar.data.pipeline 3.3.0
Major Changes
- All summary tables (taxa, districts, gear) now include a ‘date’ (monthly) column and are grouped by month. Missing months are filled as NA for consistent time series exports.
peskas.zanzibar.data.pipeline 3.1.0
peskas.zanzibar.data.pipeline 2.6.0
peskas.zanzibar.data.pipeline 2.5.0
Major Changes
- Enhanced validation workflow with KoboToolbox integration:
- Added
update_validation_status()function to update submission status via API - Added
sync_validation_submissions()for parallel processing of validation flags - Updated Kobo URL endpoint from kf.kobotoolbox.org to eu.kobotoolbox.org
- Added
New Features
- Implemented parallel processing for validation operations using future/furrr packages
- Added progress reporting during validation operations via progressr package
- Enhanced validation status synchronization between local system and KoboToolbox
Improvements
- Updated data preprocessing to handle flying fish estimates and taxa corrections (TUN→TUS, SKH→CVX)
- Updated export workflow to use validation status instead of flags for data filtering
- Added taxa information to catch export data
- Added Zanzibar SSF report template with visualization examples
- Improved package documentation structure with better categorization
peskas.zanzibar.data.pipeline 2.4.0
Major Changes
- Implemented support for multiple survey data sources:
- Refactored
get_validated_surveys()to handle WCS, WF, and BA sources - Added source parameter to specify which datasets to retrieve
- Improved handling of data sources with different column structures
- Refactored
New Features
- Added
export_wf_data()function for WorldFish-specific data export - Enhanced validation with additional composite metrics:
- Price per kg validation
- CPUE (Catch Per Unit Effort) validation
- RPUE (Revenue Per Unit Effort) validation
Improvements
- Added min_length parameter for better length validation thresholds
- Updated LW coefficient filtering logic in model-taxa.R
- Enhanced alert flag handling with combined flags from different validation steps
- Improved catch price and catch weight handling for zero-catch outcomes
- Enhanced data preprocessing with better field type conversion
peskas.zanzibar.data.pipeline 2.3.0
Major Changes
- Enhanced KoboToolbox integration:
- Implemented new validation status retrieval from KoboToolbox API
- Updated validation workflow to incorporate submission validation status
- Improved data validation process through direct API integration
New Features
- New KoboToolbox interaction functions:
-
get_validation_status(): Retrieves submission validation status from KoboToolbox API
-
peskas.zanzibar.data.pipeline 2.2.0
Major Changes
- Completely restructured taxonomic data processing:
- Introduced new modular functions for taxa handling in model-taxa.R
- Added efficient batch processing for species matching
- Implemented optimized FAO area retrieval system
- Streamlined length-weight coefficient calculations
- Enhanced integration with FishBase and SeaLifeBase
New Features
- New taxonomic processing functions:
-
load_taxa_databases(): Unified database loading from FishBase and SeaLifeBase -
process_species_list(): Enhanced species list processing with taxonomic ranks -
match_species_from_taxa(): Improved species matching across databases -
get_species_areas_batch(): Efficient FAO area retrieval -
get_length_weight_batch(): Optimized length-weight parameter retrieval
-
Improvements
- Enhanced performance through batch processing
- Reduced API calls to external databases
- Better error handling and input validation
- More comprehensive documentation
- Improved code organization and modularity
peskas.zanzibar.data.pipeline 2.1.0
Major Changes
- Enhanced taxonomic and catch data processing capabilities:
- Added comprehensive functions for species and catch data processing
- Implemented length-weight coefficient retrieval from FishBase and SeaLifeBase
- Created functions for calculating catch weights using multiple methods
- Added new data reshaping utilities for species and catch information
- Extended Wild Fishing (WF) survey validation with detailed quality checks
- Updated cloud storage and data download/upload functions
peskas.zanzibar.data.pipeline 2.0.0
Major Changes
- Complete overhaul of the data pipeline architecture
- Added PDS (Pelagic Data Systems) integration:
- New trip ingestion and preprocessing functionality
- GPS track data processing capabilities
- Implemented MongoDB export and storage functions
- Removed renv dependency management for improved reliability
- Updated Docker configuration for more robust builds
peskas.zanzibar.data.pipeline 1.0.0
peskas.zanzibar.data.pipeline 0.2.0
New features
Added the validation step and updated the preprocessing step for wcs kobo surveys data, see preprocess_wcs_surveys() and validate_wcs_surveys() functions. Currently, validation for catch weight, length and market values are obtained using median absolute deviation method (MAD) leveraging on the k parameters of the univOutl::LocScaleB function.
In order to accurately spot any outliers, validation is performed based on gear type and species.
N.B. VALIDATION PARAMETERS ARE NOT YET TUNED
peskas.zanzibar.data.pipeline 0.1.0
Drop parent repository code (peskas.timor.pipeline), add infrastructure to download WCS survey data and upload it to cloud storage providers
New features
- The ingestion of WCS Zanzibar surveys is implemented in
ingest_wcs_surveys(). - The functions
retrieve_wcs_surveys()downloads WCS Zanzibar surveys data