Skip to contents

peskas.zanzibar.data.pipeline 4.15.0

A second estimate of total catch and revenue, by FAO’s ARTFISH method

  • NEW Monthly total catch and revenue are also estimated with the FAO ARTFISH method for each district, split by gear or boat type, beside the GPS tracker method’s estimate so the two can be compared.
  • FIXED Estimated catch and revenue are no longer shown for a district and month with fewer than ten surveyed trips, where a single trip stood for the whole fleet.

peskas.zanzibar.data.pipeline 4.14.0

Boats are placed by where they land

  • NEW Fleet activity estimates place each tracked boat in the district where its trips land, read from its GPS track, so new and moved trackers count without being linked to a district by hand.
  • NEW Fishing trips longer than two days now count in the fleet activity estimates.

peskas.zanzibar.data.pipeline 4.13.1

Reviewers’ decisions are kept between runs

  • FIXED Surveys a reviewer approves in the Peskas Management Platform stay in the data (5 on 2026-09-28), those a reviewer rejects leave it, and reviewers’ decisions are no longer undone by the next run.

peskas.zanzibar.data.pipeline 4.13.0

Records say which organization collected them

  • NEW

survey_organization names the organization behind each record, as the first column — "ZAFIRI" here. survey_id identifies the form, not the organization, and a country can run more than one programme at once, as Kenya does.

peskas.zanzibar.data.pipeline 4.12.0

A length with no species bound is now bounded anyway

  • FIXED

Alert 4 compares a catch length against a per-taxon bound from FishBase, and a taxon FishBase cannot resolve had no bound at all — so a missing bound passed silently. That covered 21.6% of rows carrying a length, including every MZZ row, and is how a 130,000 cm fish reached the API. An absolute 500 cm ceiling now backs the per-taxon rule up.

peskas.zanzibar.data.pipeline 4.11.0

Unidentified catch gets the ASFIS code for unidentified catch

The WF form offers two “not identified to species” buckets, MZZ (miscellaneous) and UNK (unknown). Only the first was mapped. A catch recorded under UNK carried no species code, so alpha3_code stayed NULL all the way to the API, and its weight could not be attributed to anything.

peskas.zanzibar.data.pipeline 4.10.0

Fixes from a cross-country audit of the validated landings parquet every country publishes to gs://peskas-api-prod. The invariant under audit — tot_catch_kg == sum(catch_kg) within trip_id at the moment of export — held at 0% failure before these changes and still holds at 0% after, in both exports. Measured against wf-surveys-validated__20260913024858_f4f1dd9__.parquet (16,711 rows, 10,862 trips) and wf-surveys-preprocessed__20260913024052_f4f1dd9__.parquet (24,036 rows, 13,857 trips): the validated export goes to 16,710 rows / 10,861 trips with the total catch unchanged at 879,764.6 kg, and the raw export keeps its row and trip counts.

Bug Fixes

  • n_fishers = 0 now raises alert 13 instead of being published (validate_wf_surveys()): the API schema declares n_fishers with a minimum of 1, but a trip whose three fisher counts were all zero was published as 0, and every per-fisher metric divided by it into Inf or NaN. Zero fishers is not a count — the trip happened and the crew was simply never entered. The WF validator had no check for it at all: validate_wcs_surveys() in the same file already flags n_fishers <= 0 as alert 13 and infinite effort indicators as alert 12, and validate_ba_surveys() flags a zero crew as its alert 2, but the WF branch carried neither. Zero-crew trips were therefore excluded only by accident, when cpue/rpue divided into Inf and tripped the outlier alerts 9 and 10 — which catches a zero crew that landed a catch and misses one that did not, because 0 / 0 is NaN, not Inf.

    validate_wf_surveys() now mirrors the WCS branch exactly: alerts 9 and 10 are guarded with !is.infinite(), alert 12 fires on an infinite effort indicator, and alert 13 fires on n_fishers <= 0. Alert 13 is deliberately not narrowed by catch_outcome — that narrowing is what let the same defect through on no-catch trips. (Mozambique’s validate_surveys_adnap() carries the identical fix as its alert 11; the code number differs because Zanzibar already uses 11 for the landing-date check and 13 for this condition in its WCS branch.)

    Data impact: one submission leaves the validated set — 646533624 (2025-02-24, Hand Line), caught by alert 13 alone, since its catch_outcome is "0" and 0 / 0 gives NaN rather than Inf. Published rows go 16,711 → 16,710 and trips 10,862 → 10,861, removing the only n_fishers = 0 row; the published minimum goes from 0 to 1. No catch is lost — the submission records catch_kg = 0 — and the total stays 879,764.6 kg. The three other zero-crew submissions (695437720, 725342728, 763307846, reporting 2, 77 and 375 kg) were already excluded and remain so, now attributed to alerts 12 and 13 rather than misreported as CPUE/RPUE outliers. The !is.infinite() guards cannot readmit anything: every infinite indicator in the current data comes from a zero crew, no submission has trip_duration = 0, and alert 12 re-catches any that alerts 9/10 no longer raise. None of these submissions is discarded — they surface in the validation flags collection for the enumerator to correct at source.

    The raw export is deliberately not covered. export_api_raw() reads the preprocessed parquet, which validation never rewrites, so it still publishes all four zero-crew submissions — 5 rows across 4 trips, carrying 0, 2, 77 and 375 kg. That is the two-stage export working as designed: the raw stage is pre-validation by definition, and Mozambique’s fix scopes to the validated set the same way. Flagged rather than fixed, because suppressing rows there would make the raw export no longer raw; if the API schema’s minimum: 1 is meant to bind the raw feed too, that is a separate decision.

  • catch_habitat now reads "Open sea", not "Open Sea" (preprocess_wf_surveys()): Kenya and Mozambique both write "Open sea" for the same concept, so Zanzibar’s casing split the category into two groups in every cross-country GROUP BY. Zanzibar was the odd one out, so Zanzibar changed. Data impact: validated export, 3,169 of 16,710 rows (19%); raw export, 4,441 of 24,036 rows (18%). Only that one column changes, the rest of the habitat vocabulary is untouched, and NA still passes through as NA. Both exports are affected because the mapping lives upstream of the validated parquet, in preprocessing — so this takes effect on the next preprocessing run, not retroactively.

Robustness

  • tot_catch_kg is now derived after deduplication, not before (format_api_wf(), format_api_wcs()): sum(catch_kg) was computed inside group_by(trip_id) while distinct() ran at the end of the pipeline, so any duplicate row reaching that point would be counted in the total and then dropped from the output — silently breaking tot_catch_kg == sum(catch_kg) for that trip. Moving distinct() ahead of the grouped sum ties the total to the rows actually published. Applied to both format_api_wf() and format_api_wcs(), and to both the raw and validated exports through them. Data impact on current data: none — Zanzibar carries no duplicates at that point, row and trip counts are unchanged by this change in isolation, and the output is all.equal()-identical. This is preventative: it is the same ordering that produced a 24.4% invariant failure rate in the Kenya pipeline. Injecting 138 duplicate rows across 100 trips breaks 98 trips (0.90%) under the old ordering with exactly doubled totals, and 0 under the new one.

Refactor

  • WCS is now excluded from the API export by configuration rather than by a silent filter (export_api_raw(), export_api_validated()): both functions downloaded the WCS survey data, ran the whole format_api_wcs() transform over it, bound the result, and then dropped every one of those rows again with dplyr::filter(!survey_id == conf$ingestion$wcs$asset_id). The published dataset never contained a WCS row, so this cost a cloud download and a full transform per run for nothing, and the reason was recorded nowhere in the code.

    The exclusion is temporary, not permanent, and the filter has been replaced with an api$include_wcs flag (default false) rather than deleted. WCS is not a retired programme: it is still ingested, preprocessed and validated on every scheduled run, and its validated file currently holds 51,296 rows across 27,472 submissions running to the present day — three times the WF volume. The filter was introduced in the same commit as format_api_wcs() itself (3260cb6, “Feat wcs”), on the same day issue #4 — “Investigate systematic differences in catch and price values between WCS and WF surveys before joint production use” — was opened. It was the hold put in place while that investigation ran.

    That audit concluded the WF/WCS gap is real rather than an artefact — both programmes sample valid operations, but in different proportions and from different vessel platforms — and left two unmet preconditions for publishing them together: every row needs a source label, and catch_price denotes different quantities in each programme (WCS derives it from market medians, WF reads a trip-level field). Those are unresolved, so the hold stands, stated in inst/config.yml, in both functions’ documentation, and in a log line on every run. format_api_wcs() is retained and reachable: flipping the flag restores the merged export.

    This is also why Zanzibar’s catch_price is 100% null in production — only the WF branch survives, and it sets catch_price = NA_real_ while mapping the form’s trip-level price to tot_catch_price. Resolving that is the catch_price semantics question issue #4 left open, not a casualty of this change.

peskas.zanzibar.data.pipeline 4.9.1

Refactor

  • The total-length restatement now comes from coasts: get_length_conversions() and convert_lw_to_tl() are deleted; get_length_weight_batch() calls coasts::convert_lw_to_tl() instead. The same arithmetic existed here, in Mozambique and in Timor, written three times with different semantics. coasts passes unconvertible rows through rather than dropping them, so the call site keeps Type == "TL" to preserve the behaviour this pipeline had. Verified against the live taxa list at FishBase 25.04 / SeaLifeBase 24.07, area 51: the same 16 rows restated for the same 7 taxa (BET, BLM, BUM, MLS, NXT, QJR, SWO), and the final table is identical() across all 43 codes — zero change to any published coefficient.
  • coasts (>= 4.13.0) is now a declared floor in DESCRIPTION.

peskas.zanzibar.data.pipeline 4.9.0

Bug Fixes

  • FishBase releases are now pinned, and a missing coefficient fails the run: rfishbase reads a remote parquet dataset over the network, so an unpinned "latest" let a new FishBase release reach the pipeline the moment a container was rebuilt, with no code change. Release 26.06 dissolved Caesionidae into Lutjanidae and Scaridae into Labridae — both family names survive with zero species in them — so any taxon named after one expanded to nothing, got no length-weight coefficients, and weighed NA, which sums to zero. The releases are now pinned per server in inst/config.yml under metadata:fishbase and threaded through all five rfishbase reads in getLWCoeffs(), which previously could mix snapshots within a single run. rfishbase is additionally pinned to 5.0.1 in both Dockerfiles as the last install step, because remotes::install_local(dependencies = TRUE) upgrades it otherwise.

  • assert_taxa_coverage() fails the run when a taxon resolves to no coefficients: previously a taxon that matched nothing was dropped in silence and the pipeline stayed green while publishing a hole.

  • Unmatched taxa are now logged: match_species_from_taxa() dropped any name that matched no species without a warning.

  • SeaLifeBase routing corrected: process_species_list() routed only ISSCAAP groups 57, 45, 43, 42 and 56 to SeaLifeBase, sending sea cucumbers, gastropods, oysters, mussels, scallops and mantis shrimp to FishBase, where they matched nothing and were dropped. Routing is now ISSCAAP >= 40.

  • Species names ending in “idae” are no longer read as families: the rank test placed the family suffix before the species test, so Haliotis midae and Jordanella floridae were searched as families and matched nothing.

  • Length-type conversion recovers 16 taxa that weighed NA (get_length_conversions(), convert_lw_to_tl()): FishBase tags every published length-weight pair with the length type the original study measured, and for tunas, billfish and several carangids that is fork length. get_length_weight_batch() kept only Type == "TL", so those taxa got no coefficients at all and every length-measured catch row of them weighed NA. The conversions are published data in FishBase’s POPLL table, which this pipeline never read; coasts::get_taxa_morphometrics() already does. Reading it restates a on a total-length basis (a_TL = a * ratio^b, b unchanged) and recovers ALB BET BLM BUM CJC EWM FLY LJK LTQ LWO MLS NAB NXM NXP QJR SWO. Coverage goes from 98 to 114 of 134 codes, and none of the 98 coefficients that already resolved changed — the conversion is applied only to taxa that would otherwise have nothing.

    Validated against the 632 species carrying both a native TL pair and an FL one: converting halves the median error in predicted weight (15.6% against 27.5% for using the FL pair as-is), and against the 1,186 with both TL and SL pairs it cuts it fourfold (16.8% against 70.6%). Ratios are physically sensible (median FL/TL 0.962, SL/TL 0.831). Compared against FishBase’s independent TL-basis Bayesian estimates for the recovered taxa, converting is closer than raw on 10 of 15.

  • Retired survey codes are remapped on read (reshape_catch_data()): editing a KoBoToolbox form’s choice list does not rewrite submissions already collected, so a retired code persists in historical data indefinitely. AHI (171 rows) and BFL (6 rows) had additionally been removed from the Airtable taxa table, which orphaned them — map_surveys() left those rows with NA scientific name and alpha3 code. They now remap to BAF (Ablennes hians) and TEI (Pterocaesio pisang), joining the TUN and SKH remaps that were already there.

  • The MAC correction now runs early enough to matter: MAC was wrongly offered under the sharks-and-rays group in the form, and all 150 rows carrying it are fish_group == "SR" with a median length of 85 cm — eagle rays, not the Atlantic mackerel ASFIS maps MAC to. The SR/MAC → AQX rule existed in preprocess_wf_surveys() but ran after calculate_catch(), so the rows were weighed as MAC (which resolves to no coefficients, giving NA) and only relabelled afterwards. Moving it into reshape_catch_data() lets them pick up the AQX coefficients: 137 of 190 AQX rows now carry a weight, median 34.9 kg, where the 150 previously carried none.

  • Morphology bounds no longer silently disable length validation: min(CommonLength, na.rm = TRUE) returns Inf for a taxon whose matched species all lack that field, and the permissiveness step then computed Inf - 0.75 * Inf = NaN. Every comparison against NaN is NA, which case_when() treats as no-match, so alert codes 3 and 4 never fired for those taxa — a missing bound was indistinguishable from a passed check. safe_min() now yields NA rather than Inf, and because FishBase populates CommonLength for only 10% of species against 91% for Length, missing values are estimated as 0.625 * Length — the median ratio across the 3,748 species carrying both (common_length_ratio()). All 128 taxa with morphology now have usable bounds, against 20 previously broken. Expect a wave of new length alerts on the first run: those records were never checked before.

  • Search-name aliases fix six taxa the ASFIS names could not match (taxa_search_aliases(), apply_taxa_aliases()): a handful of ASFIS reference names match nothing in the taxonomic backbone, so the taxon is dropped and every catch row of it weighs NA. CLP is named Clupeidae, which FishBase emptied of Indo-Pacific species in 2022 — it now searches Dorosomatidae, matching the row already in Timor’s equivalent table. The rest are synonyms that have moved on: ESR to Stolephorus commersonnii, RPO to Parupeneus macronemus, LZV to Ellochelon vaigiensis, OQC to Octopus cyanea, and VMX — Valamugil, a genus the backbone no longer carries — to Osteomugil and Moolgarda. The table carries an explicit rank because the suffix rules in process_species_list() cannot recognise a bare genus name.

    CRA (“marine crabs nei”, Brachyura) is deliberately not aliased: it is an infraorder, and SeaLifeBase carries no rank between order Decapoda and family, so choosing a target means deciding which crab families Zanzibar lands.

  • OQC now gets the octopus mantle-length conversion: OCZ (Octopus spp) was special-cased in three places — an ML-only coefficient filter, the arm-span-to-mantle /5.5 conversion in calculate_catch(), and the min_length floor — but OQC (Octopus cyaneus) was not, despite being the larger of the two in the data at 694 rows with 484 lengths, median 85 cm. OQC resolves to a mantle-length pair, so applying it to arm-span unconverted weighed a single octopus at 263 kg instead of 2.43 kg. This was latent while OQC had no coefficients and would have gone live with the alias above.

  • Fixed a duplicate join key for FLY: preprocess_wf_surveys() appended a hardcoded flying-fish coefficient unconditionally. Now that the conversion recovers Exocoetidae pairs, getLWCoeffs() returns a FLY row of its own, and two rows on the same key would have doubled every flying fish catch record in calculate_catch(). The manual value now replaces rather than appends, and stays authoritative: FishBase’s recovered pairs weigh a 30 cm flying fish at 498 g against the hardcoded 202 g, and changing that is a separate decision.

  • Removed a dead fallback in preprocess_wf_surveys(): the tryCatch around getLWCoeffs() read inst/length_weight_params.rds, which is not in the package, so the fallback could only ever fail — while hiding the original error behind it.

Known Issues

Measured 2026-09-06 against FishBase 25.04 / SeaLifeBase 24.07 over the live KoBo data: 122 of 134 codes resolve length-weight coefficients, up from 98 before this release. The other 12 form the documented baseline in assert_taxa_coverage(), so any new loss fails the run. CJX and PWT are deliberately not in it — they resolve at 25.04 and are the two codes that break at 26.06, so a release move fails the check.

  • Wrong reference name (2) — MAE, TAG name species absent from FAO 51.
  • No published coefficients (1) — GQT (Plectorhinchus gaterinus) occurs in FAO 51 but FishBase carries no length-weight pair for it in any length type. Nothing to convert, nothing to alias.
  • No convertible length type (5) — KAK, LHV, RMB, RTY, SSP.
  • Not a taxon, or a rank the backbone omits (4) — MZZ, UNKN, UNK, and CRA (the infraorder Brachyura).

peskas.zanzibar.data.pipeline 4.8.0

New Features

  • Gleaning Survey Data Integration: Added support for KoBoToolbox gleaning survey data collection and processing
    • New survey type (gleaning) is now ingested from KoBoToolbox, preprocessed, and validated alongside existing catch surveys
    • Gleaning data captures informal fisheries activity not covered by formal catch surveys, providing a more complete picture of small-scale fishing effort
    • Gleaning submissions are now included in the unified validation pipeline with tailored quality checks

peskas.zanzibar.data.pipeline 4.7.0

New Features

  • Unified API export across survey programs: The raw and validated API exports now combine data from both WCS and WorldFish surveys into a single output file. Previously only WorldFish data was exported; WCS trips are now included alongside them with a consistent set of fields.

Improvements

  • WCS validation enhancements:
    • Added two new quality checks: one that catches contradictions between bucket count and bucket weight (e.g. buckets recorded but no weight, or vice versa), and one that flags implausible negative values in catch measurements
    • Price and revenue validation thresholds corrected to Tanzanian Shilling values (previous values were in Mozambican metical)
    • Submissions recorded before 2020 are now excluded from the validated dataset
    • Catch outcome is now correctly carried through to the validation step
  • WCS price calculation corrected: Catch prices are now split proportionally across species within a trip. Missing species prices now fall back to the species-level median rather than being left empty.

peskas.zanzibar.data.pipeline 4.6.0

Infrastructure & Workflow

  • Delegating to coasts most of the core storage and databse-related functions: Now core and other countries shared storage functions are delagated to central and upgraded features of the coastspacakge for improved standardization and maintainability

peskas.zanzibar.data.pipeline 4.5.0

Major Changes

  • Adopted coasts as the shared multicountry analytics engine: Aggregated data summarization and dashboard export are now delegated to WorldFishCenter/peskas.coasts (dev branch). This centralizes the logic for producing monthly, taxa, district, and gear summaries — as well as fishery metrics — across all Peskas country deployments (Zanzibar, Kenya, Mozambique), ensuring consistent outputs and a single place to maintain and improve the shared pipeline logic.
    • Added coasts to Imports and Remotes in DESCRIPTION
    • Added remotes::install_github("WorldFishCenter/peskas.coasts", ref = "dev") to both Dockerfile and Dockerfile.prod so the image ships the package
    • Pipeline steps that previously used local summarize_data() and generate_fleet_analysis() now call the equivalent coasts:: functions, passing package = "peskas.zanzibar.data.pipeline" so they read the country-specific inst/conf.yml

peskas.zanzibar.data.pipeline 4.4.0

Improvements

  • Standardized configuration structure: Replaced inst/conf.yml with a unified multi-country template harmonized across all Peskas deployments (Zanzibar, Kenya, Mozambique). Key structural changes:
    • Survey credentials moved from surveys.* into a new top-level ingestion.* section
    • Stage keys shortened (raw_surveys → raw, preprocessed_surveys → preprocessed, etc.)
    • Source names shortened (wcs_surveys → wcs, wf_surveys_v1 → wf_v1, etc.)
    • MongoDB structure reorganized: connection strings under connection_strings.*, databases under databases.*, collections key pluralized, portal renamed to dashboard
    • Airtable config moved from top-level airtable.* to metadata.airtable.*
    • All R code updated to use the new config paths
  • summarize_data() correctness and clarity fixes:
    • Fixed data quality bug: taxa_summaries was incorrectly summing trip-total catch kg per taxon instead of the actual per-taxon catch weight — values are now correct
    • Fixed districts_summaries round-trip pivot: data is now stored wide and pivoted once in export_wf_data() instead of pivot-long → store → pivot-wide → pivot-long
    • Fixed gear_summaries complete() running after pivot_longer with wrong column names in the fill list; now completes before pivoting across all gear × district × month combinations
    • Replaced fragile across(everything(), first) pattern in trip-level collapse with explicit slice(1), which is clearer and avoids silent column overwrites
    • Fixed inconsistent na.rm usage in gear summaries
  • Removed dead code: Deleted create_geos() and create_geos_v1() from export.R; the geographic summary logic had already been inlined into export_wf_data()

Bug Fixes

  • Fixed match_surveys_to_registry.Rd cross-reference warning caused by [0,1] being parsed as a markdown link
  • Documented missing devices_table argument in process_trip_data()
  • Added stringdist to Imports in DESCRIPTION (was used via :: but not declared)

peskas.zanzibar.data.pipeline 4.3.0

New Features

  • Survey-GPS Trip Matching Pipeline: Added comprehensive fuzzy matching system to link catch survey records with GPS trip data
    • New merge_trips() workflow function for Kenya and Zanzibar sites
    • match_surveys_to_gps_trips(): Universal two-step matching (surveys → registry → trips)
    • Fuzzy string matching using Levenshtein distance on registration numbers, boat names, and fisher names
    • Conservative one-trip-per-day constraint to avoid ambiguous matches
    • Support for both explicit device registries (Kenya) and implicit registries built from trip data (Zanzibar)
    • Configurable matching thresholds (registration: 15%, names: 25% difference allowed)
    • Exports merged dataset with matched pairs plus all unmatched surveys and trips
    • Helper functions: standardize_column_names(), clean_matching_fields(), build_registry_from_trips()

Improvements

  • PDS Data Ingestion:
    • Updated ingest_pds_trips() to load device registry from cloud storage instead of Airtable metadata
    • Added device info retrieval and proper filtering for Zanzibar devices
    • Improved configuration variable naming (pars → conf)
  • GitHub Actions Workflow:
    • Added new merge-trips job to automated pipeline
    • Runs after survey preprocessing and before summarization
    • Integrated with production environment configuration

peskas.zanzibar.data.pipeline 4.2.0

New Features

  • API Data Export Pipeline: Added new export_api_raw() function to export raw preprocessed survey data in API-friendly format
    • Exports raw/preprocessed trip data (before validation) to cloud storage
    • Part of a two-stage API export pipeline (raw and validated exports)
    • Transforms nested survey data into flat structure with standardized trip-level records
    • Generates unique trip IDs using xxhash64 algorithm
    • Integrates with Airtable metadata for form-specific asset lookups
    • Exports versioned parquet files to zanzibar/raw/ path for external API consumption
    • Includes comprehensive output schema with 14 standardized fields (trip_id, landing_date, gear, catch metrics, etc.)
  • Airtable Integration: New helper functions for managing Airtable metadata and form configurations

Improvements

  • Configuration Enhancements:
    • Added api configuration section for trip data exports with separate raw/validated paths
    • Configured cloud storage paths for API exports (zanzibar/raw, zanzibar/validated)
    • Added Airtable base ID and token configuration for metadata management
    • Enhanced options_api storage configuration for peskas-coasts bucket
  • GitHub Actions Workflow:
    • Added new export-api-data job to automated pipeline workflow
    • Integrated Airtable authentication with GitHub Secrets (AIRTABLE_TOKEN, AIRTABLE_BASE_ID_FRAME, AIRTABLE_BASE_ID_ASSETS)
    • Configured API export job to run after survey preprocessing step
    • Added production environment configuration for API data exports
  • Code Quality:
    • Improved documentation with comprehensive roxygen2 comments for all new functions
    • Added detailed examples and cross-references in function documentation
    • Enhanced error handling and input validation in Airtable operations
    • Implemented proper cleanup of temporary local files after cloud uploads

peskas.zanzibar.data.pipeline 4.1.1

Major Changes

  • Streamlined Validation Workflow: Replaced KoboToolbox API updates with direct MongoDB storage to improve performance.
    • New export_validation_flags() function exports validation flags directly to MongoDB
    • Validation status queries now only identify manually edited submissions, not update them
    • Disabled sync_validation_submissions() workflow steps in GitHub Actions
    • Significantly reduced pipeline execution time by avoiding slow KoboToolbox API calls

Improvements

  • Validation System:

    • Validation functions now preserve manual human approvals while updating system-generated statuses
  • Code Quality:

    • Fixed SeaLifeBase API calls by pinning to version 24.07 to avoid server errors
    • Standardized function parameter formatting across validation and preprocessing modules

peskas.zanzibar.data.pipeline 4.1.0

Major Changes

  • Integration of New KoBoToolbox Survey Form Version: Added support for a new version of the WorldFish survey form (wf_surveys_v2) alongside the existing form (wf_surveys_v1). Data from both survey versions is now processed together in the preprocessing pipeline and handled properly throughout the validation workflow.

Improvements

  • Multi-Asset Validation Support:
    • Updated validation system to query approval statuses from both survey form versions
    • Enhanced validate_wf_surveys() and sync_validation_submissions() to handle submissions from multiple KoBoToolbox assets
    • Ensured manually approved submissions from either form version are protected from automated flagging
  • Configuration Updates:
    • Added configuration for the new survey form version with shared credentials
    • Cleaned up redundant configuration entries
    • Updated code references to use versioned asset configurations

Bug Fixes

  • Fixed validation logic that was only checking approval status from the original survey form, causing incorrect flagging of valid submissions from the new form version

peskas.zanzibar.data.pipeline 4.0.0

Major Changes

  • Fleet Activity Analysis Pipeline: Introduced a comprehensive pipeline for estimating and analyzing fishing fleet activity using GPS-tracked boats and boat registry data. This includes new functions for preparing boat registries, processing trip data, calculating monthly trip statistics, estimating fleet-wide activity, and calculating district-level total catch and revenue.
  • New Modeling and Summarization Functions:
    • prepare_boat_registry(): Summarizes boat registry data by district.
    • process_trip_data(): Processes trip data with district information and filters outliers.
    • calculate_monthly_trip_stats(): Computes monthly fishing activity statistics by district.
    • estimate_fleet_activity(): Scales up sample-based trip statistics to fleet-wide estimates.
    • calculate_district_totals(): Combines fleet activity and catch data for district-level totals.
    • generate_fleet_analysis(): Orchestrates the full analysis pipeline and uploads results.
    • summarize_data(): Generates and uploads summary datasets (monthly, taxa, district, gear, grid) for WorldFish survey data.
  • Enhanced Data Export and Integration:
    • export_wf_data(): Exports summarized WorldFish survey data and modeled estimates to MongoDB, including new geographic regional summaries.
    • create_geos(): Generates geospatial regional summaries and exports as GeoJSON for spatial visualization.
  • Expanded Documentation: New and updated Rd files for all major new functions, with improved examples and cross-references.

Improvements

  • Consistent Time Series and Grouping: All summary tables (taxa, districts, gear) now include a ‘date’ (monthly) column and are grouped by month, with missing months filled as NA for consistent time series exports.
  • Parallel Processing: Improved use of parallelization (via future and furrr) for validation and summarization steps, enhancing performance for large datasets.
  • Data Quality and Validation:
    • Enhanced filtering and validation of survey data before summarization and export.
    • Improved handling of flagged/invalid submissions.

Infrastructure & Workflow

  • Configuration and Documentation: Updated configuration files and documentation to support new modeling and export workflows.
  • Workflow Automation: Updates to GitHub Actions and Docker configuration to support the expanded pipeline.

peskas.zanzibar.data.pipeline 3.3.0

Major Changes

  • All summary tables (taxa, districts, gear) now include a ‘date’ (monthly) column and are grouped by month. Missing months are filled as NA for consistent time series exports.

Validation Updates

  • The maximum number of individuals per catch is now 200 (was 80).
  • Validation flag for ‘number of fishers too high’ is now triggered at >100 (was >70) for non-ring nets.
  • Documentation updated to reflect new validation thresholds.

Code Quality

  • Improved code formatting and clarity in validation functions and documentation.

peskas.zanzibar.data.pipeline 3.2.0

Improvements

  • Export standardized fishery metrics for general usage with other peskas datasets

peskas.zanzibar.data.pipeline 3.1.0

New Features

  • Added create_geos() function to generate geospatial regional summaries of fishery data
  • Added support for GPS track data visualization through new grid-based analytics
  • Added generate_track_summaries() function to process GPS tracks into 1km grid cells

Improvements

  • Integrated spatial data with dashboard exports through new GeoJSON support
  • Enhanced code readability and formatting throughout codebase
  • Added Region-based aggregation of fishery metrics (CPUE, RPUE, price per kg)
  • Added grid summaries to MongoDB exports for dashboard integration

peskas.zanzibar.data.pipeline 3.0.0

New Features

  • Generate clean dataframes to export to dashboard. These include districts and taxa summaries and monthly regional time series of the main fishery indicators, CPUE, RPUE and price per kg

peskas.zanzibar.data.pipeline 2.6.0

New Features

  • Added new sync-validation job to GitHub Actions workflow for synchronizing survey validation submissions

Improvements

  • Implemented error handling in getLWCoeffs to fallback on local data if Rfishbase retrieval fails
  • Enhanced code readability by restructuring functions and adding line breaks
  • Updated documentation for get_preprocessed_surveys and get_validated_surveys functions

peskas.zanzibar.data.pipeline 2.5.0

Major Changes

  • Enhanced validation workflow with KoboToolbox integration:
    • Added update_validation_status() function to update submission status via API
    • Added sync_validation_submissions() for parallel processing of validation flags
    • Updated Kobo URL endpoint from kf.kobotoolbox.org to eu.kobotoolbox.org

New Features

  • Implemented parallel processing for validation operations using future/furrr packages
  • Added progress reporting during validation operations via progressr package
  • Enhanced validation status synchronization between local system and KoboToolbox

Improvements

  • Updated data preprocessing to handle flying fish estimates and taxa corrections (TUN→TUS, SKH→CVX)
  • Updated export workflow to use validation status instead of flags for data filtering
  • Added taxa information to catch export data
  • Added Zanzibar SSF report template with visualization examples
  • Improved package documentation structure with better categorization

peskas.zanzibar.data.pipeline 2.4.0

Major Changes

  • Implemented support for multiple survey data sources:
    • Refactored get_validated_surveys() to handle WCS, WF, and BA sources
    • Added source parameter to specify which datasets to retrieve
    • Improved handling of data sources with different column structures

New Features

  • Added export_wf_data() function for WorldFish-specific data export
  • Enhanced validation with additional composite metrics:
    • Price per kg validation
    • CPUE (Catch Per Unit Effort) validation
    • RPUE (Revenue Per Unit Effort) validation

Improvements

  • Added min_length parameter for better length validation thresholds
  • Updated LW coefficient filtering logic in model-taxa.R
  • Enhanced alert flag handling with combined flags from different validation steps
  • Improved catch price and catch weight handling for zero-catch outcomes
  • Enhanced data preprocessing with better field type conversion

Bug Fixes

  • Fixed issue with catch_price field type in WF survey preprocessing
  • Corrected filter condition for taxa coefficients

peskas.zanzibar.data.pipeline 2.3.0

Major Changes

  • Enhanced KoboToolbox integration:
    • Implemented new validation status retrieval from KoboToolbox API
    • Updated validation workflow to incorporate submission validation status
    • Improved data validation process through direct API integration

New Features

  • New KoboToolbox interaction functions:
    • get_validation_status(): Retrieves submission validation status from KoboToolbox API

Improvements

  • Modified configuration files to support new KoboToolbox API token
  • Added new environment variable for KoboToolbox API authentication
  • Enhanced validation workflow with integrated validation status checks

peskas.zanzibar.data.pipeline 2.2.0

Major Changes

  • Completely restructured taxonomic data processing:
    • Introduced new modular functions for taxa handling in model-taxa.R
    • Added efficient batch processing for species matching
    • Implemented optimized FAO area retrieval system
    • Streamlined length-weight coefficient calculations
    • Enhanced integration with FishBase and SeaLifeBase

New Features

Improvements

  • Enhanced performance through batch processing
  • Reduced API calls to external databases
  • Better error handling and input validation
  • More comprehensive documentation
  • Improved code organization and modularity

Deprecations

  • Removed legacy taxonomic processing functions
  • Deprecated redundant species matching methods
  • Removed outdated data transformation utilities

Documentation

  • Added detailed function documentation
  • Updated vignettes with new workflows
  • Improved code examples
  • Enhanced README with new features

peskas.zanzibar.data.pipeline 2.1.0

Major Changes
  • Enhanced taxonomic and catch data processing capabilities:
    • Added comprehensive functions for species and catch data processing
    • Implemented length-weight coefficient retrieval from FishBase and SeaLifeBase
    • Created functions for calculating catch weights using multiple methods
    • Added new data reshaping utilities for species and catch information
  • Extended Wild Fishing (WF) survey validation with detailed quality checks
  • Updated cloud storage and data download/upload functions

peskas.zanzibar.data.pipeline 2.0.0

Major Changes
  • Complete overhaul of the data pipeline architecture
  • Added PDS (Pelagic Data Systems) integration:
    • New trip ingestion and preprocessing functionality
    • GPS track data processing capabilities
  • Implemented MongoDB export and storage functions
  • Removed renv dependency management for improved reliability
  • Updated Docker configuration for more robust builds
New Features
  • Enhanced validation system for survey data
  • Added new data processing steps:
    • GPS track preprocessing
    • Catch data validation
    • Length measurements validation
    • Market data validation
  • Flexible data export capabilities
  • Improved GitHub Actions workflow with additional processing steps
Infrastructure Updates
  • Streamlined package dependencies
  • Updated build and deployment processes
  • Enhanced data storage and retrieval mechanisms

peskas.zanzibar.data.pipeline 1.0.0

Improvements
  • All the functions are now documented and indexed according to keywords
  • Thin out the R folder gathering functions by modules
Changes
  • Move to parquet format rather than CSV/RDS

peskas.zanzibar.data.pipeline 0.2.0

New features

Added the validation step and updated the preprocessing step for wcs kobo surveys data, see preprocess_wcs_surveys() and validate_wcs_surveys() functions. Currently, validation for catch weight, length and market values are obtained using median absolute deviation method (MAD) leveraging on the k parameters of the univOutl::LocScaleB function.

In order to accurately spot any outliers, validation is performed based on gear type and species.

N.B. VALIDATION PARAMETERS ARE NOT YET TUNED

Changes

No need to run the pipeline every two days, decreased not to every 4 days.

peskas.zanzibar.data.pipeline 0.1.0

Drop parent repository code (peskas.timor.pipeline), add infrastructure to download WCS survey data and upload it to cloud storage providers

New features

  • The ingestion of WCS Zanzibar surveys is implemented in ingest_wcs_surveys().
  • The functions retrieve_wcs_surveys() downloads WCS Zanzibar surveys data
Changes
  • Updated configuration management:
    • Moved configuration settings to inst/conf.yml
    • Improved configuration structure and organization
    • Enhanced configuration flexibility