Databook

activeopen government

An award-winning transparency tool that combines over 50 official NYC open datasets — agencies, people, jobs, capital projects, schools, districts and procurement — into one searchable picture of how city government works.

★ Pinned
Nothing pinned yet — set pinned on an event, knowledge item, or task in the admin.
Activity
Jul 29
PR #167: fix(scheduler): actually run the unmapped-entity scan wegovnyc/databook-private

Fixed a bug where the unmapped-entity scan never executed by moving it to a new end-of-cycle scheduler sweep. Additionally, registered the "civillist" dataset as normalized to enable its entity scanning, added six new tests, and resolved a gap identifying one unmapped agency.

@devinbalkind
Jul 29
PR #166: fix(seed): give needs_normalization a single owner wegovnyc/databook-private

This update fixes a database seeding bug by making the NORMALIZED_DATASETS membership the single owner of the needs_normalization column. It removes the redundant DIRECT_SOCRATA_OVERRIDE configuration and prevents incorrect column overwrites, resolving a 27-dataset drift issue without changing production data.

@devinbalkind
Jul 29
PR #165: fix(seed): make the registry deactivations actually run in prod wegovnyc/databook-private

This update fixes a bug where registry deactivations were bypassed in production due to a development-only path check. The deactivation logic has been moved to run unconditionally on every execution, and five new tests were added to prevent future regression without altering other production ingest behaviors.

@devinbalkind
Jul 29
PR #164: fix(registry): deactivate the second phantom CPDB dup, and guard the class wegovnyc/databook-private

Deactivated the duplicate cpdb_commitments registry entry to prevent redundant, failing normalizer runs.

@devinbalkind
Jul 29
PR #163: feat(monitoring): alert when a dataset falls behind or empties at source wegovnyc/databook-private

Added a new dataset-staleness monitoring script that alerts when Socrata datasets fall behind their source or go empty. Additionally, the import process now explicitly distinguishes empty-at-source errors from headerless responses.

@devinbalkind
Jul 29
PR #10: fix(ingest): an empty source is not a pipeline failure wegovnyc/normalizer-py

This update prevents pipeline failures when an NYC Open Data source publishes an empty dataset with a valid CSV header. Instead of raising an error, the system now logs a warning, caches the ETag to prevent redundant reprocessing, and continues serving the last good output.

@devinbalkind
Showing 8590 of 255 events
Prev15 / 43Next
Knowledge
No knowledge items yet.