Databook

activeopen government

An award-winning transparency tool that combines over 50 official NYC open datasets — agencies, people, jobs, capital projects, schools, districts and procurement — into one searchable picture of how city government works.

★ Pinned
Nothing pinned yet — set pinned on an event, knowledge item, or task in the admin.
Activity
Jun 22
PR #3: Docs: Who-Owns-What parity plan + Phase 1 api2 spec wegovnyc/map_databook

Adds planning documentation for bringing the building lookup to Who-Owns-What parity and creating a dedicated landlord portfolio view. Includes a phased implementation plan, name-normalization specifications, and SQL/endpoint specifications for Phase 1 API additions without making any code or data changes.

@devinbalkind
Jun 22
11 commits to main wegovnyc/map_databook

Added zoned-school lookups, accountability chips, and ownership sections to building profiles. Integrated a landlord portfolio view and updated documentation with Phase 1 and 2 API specifications, including Who-Owns-What parity plans and portfolio endpoint details.

@devinbalkind
Jun 18
PR #9: fix(ingest): trim leading/trailing whitespace from CSV column names wegovnyc/normalizer-py

This update fixes a bug where leading and trailing whitespace in CSV header names propagated to the API and broke the frontend. A new helper function now trims outer whitespace from column names during both network and local CSV ingestion.

@devinbalkind
Jun 18
3 commits to main wegovnyc/normalizer-py

Updated the CSV ingestion process to trim leading and trailing whitespace from column names. Fixed a bug in the match-saving process by deduping in-batch source texts to prevent duplicate database entries.

@devinbalkind
Jun 18
PR #8: fix(matches): dedupe in-batch source_texts in save_matches wegovnyc/normalizer-py

This pull request fixes a pipeline crash caused by duplicate source texts within the same batch violating a database unique constraint. It resolves the issue by deduplicating source texts in the save_matches function before the upsert loop.

@devinbalkind
Jun 16
PR #7: docs: README reflects git-driven prod deploy + /data DuckDB path wegovnyc/normalizer-py

It also corrects the DuckDB data path from a stale container directory to the active host directory.

@devinbalkind
Showing 163168 of 255 events
Prev28 / 43Next
Knowledge
No knowledge items yet.