peskas.timor.data.pipeline implements, deploys and executes the data and modelling pipelines behind Peskas, the small-scale fisheries analytics system of Timor-Leste.
It ingests KoBoToolbox landing surveys and Pelagic Data Systems (PDS) GPS tracker data, preprocesses and validates them, models fishery indicators, and publishes JSON to a public Google Cloud Storage bucket that the live portal reads. It also publishes the cross-country API tables that Peskas shares with the Kenya, Mozambique and Zanzibar pipelines, deposits datasets on Harvard Dataverse, and emails reports.
The pipeline is an R package
Structuring the pipeline as an R package makes it easier to write production-grade software. Specifically, it allows us to:
- better handle system and package dependencies,
- forces us to split the code into functions,
- makes it easier to document the code, and
- makes it easier to test the code
We make heavy use of tidyverse style conventions and the usethis package to automate tasks during project setup and deployment.
For more information about the rationale of structuring a pipeline as a package, see Chapter 3 of Engineering Production-Grade Shiny Apps. The best place to learn more about package development is the R packages book by Hadley Wickham and Jenny Bryan.
Shared code lives in peskas.coasts
Everything that is not specific to Timor-Leste has moved to peskas.coasts, the hub the four Peskas country pipelines share: cloud storage, KoBo retrieval, PDS ingestion, the Airtable reference frame, MongoDB access and the taxa morphometrics. This package keeps what is genuinely Timor’s — the survey reshaping of three form generations, sixteen validators, the glmmTMB catch and revenue models, the nutrient and RDI calculations, and the portal JSON contract.
peskas.coasts is not on CRAN. It is declared in Remotes: and, in the production container, installed from the latest tagged release, which the workflow resolves at build time and passes in as COASTS_REF. Version 4.6.0 is a hard floor.
The pipeline runs on GitHub Actions
Each step of the pipeline is a function in this package, and those functions are deployed and connected using GitHub Actions. Workflow functions take no arguments and are used for their side effects.
The pipeline is defined in .github/workflows/data-pipeline.yaml — every two days and on every push:
build-container
├── ingest-preprocess-metadata-tables
├── ingest-landings ──▶ preprocess-landings
└── ingest-pds-data ──▶ preprocess-pds-data ──▶ validate-pds-data
merge-landings ──▶ validate-landings ──▶ merge-trips ──▶ model-indicators
└─▶ export-tripsThe other workflows are R-CMD-check, pkgdown, test-coverage, pr-commands, release (which cuts a GitHub release from NEWS.md), and three scheduled Timor-specific jobs: the data report, the Dataverse upload and the weekly validation email.
Artifacts produced by each job are written to cloud storage and read back by the next job. They are versioned by add_version(), which stamps a timestamp and the commit sha, so every artifact traces to a unique run:
coasts::cloud_object_name(version = "latest") resolves the newest version of a prefix, and coasts::cloud_object_names() enumerates a whole family.
Interchange format is flat long parquet — one row per (submission, catch, length bin) — from the raw survey table through to the validated one.
Environment parameters are in the config file
How the pipeline runs is specified in inst/config.yml, read by read_config(). Its shape follows the cross-country template shipped as inst/config_template.yml. Keeping these parameters out of the code is what lets one codebase run against two sets of cloud resources; we use the config package, which selects an environment from R_CONFIG_ACTIVE.
There are two environments:
-
default— the development environment. It uses the-devbuckets (timor-dev,pds-timor-dev,public-timor-dev,peskas-api-dev), so the whole pipeline can run end to end without touching anything the portal reads. This is what.Renvironsets, and what every push to a non-mainbranch runs in CI. -
production— the same code against the production buckets. The workflow setsR_CONFIG_ACTIVE=productiononly when the ref ismain.
Credentials come from the environment in both cases, never from a file in the repository. Copy .env.example to .env and fill it in for local work — read_config() loads it through load_dotenv(). In CI the same variables are supplied from GitHub secrets by the workflow’s env: block. .env is gitignored and must stay that way.
Note that read_config() deliberately logs configuration key names only. The resolved configuration holds the service-account key and every token, and GitHub Actions masks only byte-exact matches of a registered secret.
We use docker containers
Docker makes it easier to run and develop the code.
-
Development:
Dockerfileis based onrocker/geospatialand spins up an RStudio server. Rundocker compose upfrom the project directory and open http://localhost:8802. -
Production:
Dockerfile.prodis what the pipeline runs in. The first job of the workflow builds it and every other job uses it, so each step executes in an identical environment.
Both need to know which peskas.coasts release to install, and neither has a default, so a local build must pass it explicitly:
Tests
Two suites, with different jobs:
-
tests/testthat/— unit tests over pure reshaping and schema logic. No network, no credentials. Run withdevtools::test(). -
inst/tinytest/— four assertion suites over the real artifacts a run produces (validated landings, validated PDS trips, merged trips, public data). They run as steps inside the pipeline workflow, right after the job that writes what they assert on, so a schema regression stops the pipeline instead of reaching the portal.
tinytest::run_test_file(system.file("tinytest/test_validated_landings.R",
package = "peskas.timor.data.pipeline"))The portal discovers its files dynamically, which means a renamed or dropped object silently disappears from the site. data-raw/compare-portal-json.R is the gate for that: it asserts object names, keys, nesting, column sets and column types against a reference set, and reports per-column numeric summaries. Run it before changing anything on the export path.
Logging
We use the logger package to log events in production. Every workflow function takes a log_threshold argument.