Prompt 20 · Lesson 05 · Location-service engineering

Design and Benchmark a Location Service with Precision, Bounding, and Index Selectivity

Measure radius, precision, selectivity, index cost, and latency as a complete location-service workload.

Advanced150–220 minutesPrecision/benchmark labMongoDB 8.3.8 · mongosh 2.10.0 · PyMongo 4.17.0Last reviewed: September 2026

Learning objectives

01

Design a bounded AtlasMart location-service API with explicit coordinate precision, tenant filters, maximum search radius, and result limits.

02

Generate a deterministic synthetic geographic workload and benchmark p50/p95/p99 latency without coordinated omission claims.

03

Measure query selectivity across different radii and business filters using executionStats.

04

Quantify coordinate-rounding error so storage/privacy decisions are separated from accidental precision loss.

05

Create an operational checklist for index cost, query abuse, sharding boundaries, and geospatial regression tests.

Reproducible lab baseline

This lesson pins MongoDB Community Server 8.3.8 using mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim, mongosh 2.10.0, and PyMongo 4.17.0 where client measurement is useful. The mandatory lab is a disposable loopback-only standalone on port 27169 named atlasmart-ch20-l5. Authentication and TLS are disabled only for this isolated learning container. Default read/write concern and primary read preference apply. FCV is observed and never changed. Atlas, Search, Vector Search, KMS, and Enterprise Advanced are not required. MongoDB 8.3 creates new 2dsphere indexes as index version 4; downgrading FCV below 8.3 requires dropping those version-4 indexes first. Product labs were not executed in this generation environment, so distance, explain, index-size, and latency evidence that depends on runtime must be measured locally rather than copied as invented output. The benchmark inserts 20,000 synthetic points near a fixed center with a deterministic random seed. Its runtime numbers are intentionally not pre-filled; only the workload shape is deterministic.

1. Define the API contract before measuring it

A production “find nearby pickup points” endpoint needs more than a coordinate. Define canonical GeoJSON longitude/latitude order, allowed coordinate precision, tenant/visibility predicate, maximum radius, maximum result count, whether closed locations can appear, distance units, and timeout/error behavior. Without bounds, a user can turn a local lookup into a broad geospatial workload.

Contract item AtlasMart example Why it matters
Coordinate input GeoJSON Point, longitude then latitude prevents axis ambiguity
Radius bounded meters controls candidate work and abuse
Tenant/status server-side predicate authorization + selectivity
Result limit bounded top N protects response and downstream work
Precision preserve source precision unless policy intentionally reduces it avoids accidental location drift
Evidence p50/p95/p99 + executionStats avoids tuning from averages or slogans

2. Generate a deterministic 20,000-point workload with PyMongo

start the disposable MongoDB 8.3.8 lab
docker rm -f atlasmart-ch20-l5 2>/dev/null || truedocker volume rm atlasmart-ch20-l5-data 2>/dev/null || truedocker run -d --name atlasmart-ch20-l5 \  -p 127.0.0.1:27169:27017 \  -v atlasmart-ch20-l5-data:/data/db \  mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim \  --bind_ip_alldocker exec atlasmart-ch20-l5 mongosh --quiet --eval 'printjson(db.adminCommand({buildInfo:1}).version);printjson(db.adminCommand({getParameter:1,featureCompatibilityVersion:1}).featureCompatibilityVersion);' 
load deterministic synthetic pickup locations and measure batches
from pymongo import MongoClient, GEOSPHERE, ASCENDINGfrom random import Randomfrom time import perf_counterfrom statistics import medianclient=MongoClient("mongodb://127.0.0.1:27169/?directConnection=true")coll=client.atlasmart.pickup_points_ch20_l5coll.drop()rng=Random(20260903)center_lon,center_lat=49.8671,40.4093batch=[]; batch_ms=[]for i in range(20000):    # Synthetic box roughly +/- 0.12 degrees around the center; not a real city model.    lon=center_lon+rng.uniform(-0.12,0.12)    lat=center_lat+rng.uniform(-0.12,0.12)    batch.append({"_id":f"p-{i:05d}","tenantId":"tenant-a" if i%5 else "tenant-b","open":i%4!=0,"category":"locker" if i%3==0 else "store","location":{"type":"Point","coordinates":[lon,lat]}})    if len(batch)==500:        t0=perf_counter(); coll.insert_many(batch,ordered=False); batch_ms.append((perf_counter()-t0)*1000); batch=[]coll.create_index([("tenantId",ASCENDING),("open",ASCENDING),("location",GEOSPHERE)],name="tenant_open_geo")vals=sorted(batch_ms)def pct(p):    return vals[min(len(vals)-1,round((len(vals)-1)*p))]print({"documents":coll.count_documents({}),"batches":len(vals),"p50_insert_ms":pct(.50),"p95_insert_ms":pct(.95),"p99_insert_ms":pct(.99)})client.close()

The data distribution is reproducible, but timing is not. Report CPU/RAM limits, storage, server/client versions, warm-up, concurrency, write concern, container overhead, and competing workloads with every benchmark result.

3. Measure bounding/selectivity instead of declaring one radius “fast”

compare several radii and eligibility predicates
const d=db.getSiblingDB("atlasmart");const origin={type:"Point",coordinates:[49.8671,40.4093]};for(const radiusM of [500,2000,10000,30000]){  const q={tenantId:"tenant-a",open:true,location:{$near:{$geometry:origin,$maxDistance:radiusM}}};  const t=Date.now(); const rows=d.pickup_points_ch20_l5.find(q,{_id:1}).limit(50).toArray();  printjson({radiusM,returned:rows.length,clientElapsedMs:Date.now()-t});  printjson(d.pickup_points_ch20_l5.explain("executionStats").find({tenantId:"tenant-a",open:true,location:{$geoWithin:{$centerSphere:[[49.8671,40.4093],radiusM/6371008.8]}}}));}

$near already uses $maxDistance as a useful geographic bound. The companion $geoWithin/$centerSphere explains spatial selectivity without nearest-first sorting; its radius is converted from meters to radians. Compare result counts, keys/documents examined, and elapsed distributions. A broader radius can be operationally more expensive even when the same index is used.

4. Quantify coordinate precision loss

calculate drift introduced by decimal rounding
from math import radians,sin,cos,asin,sqrtdef h(a,b):    lon1,lat1=map(radians,a); lon2,lat2=map(radians,b)    q=sin((lat2-lat1)/2)**2+cos(lat1)*cos(lat2)*sin((lon2-lon1)/2)**2    return 2*6371008.8*asin(sqrt(q))p=(49.8671347,40.4093128)for digits in [2,3,4,5,6]:    r=(round(p[0],digits),round(p[1],digits))    print({"digits":digits,"rounded":r,"drift_m":round(h(p,r),3)})

The numbers from this calculation are deterministic for the fixture and help turn “precision” into an engineering choice. Do not choose decimal truncation casually: location privacy, routing accuracy, legal zones, address resolution, and index selectivity are different requirements. If you intentionally reduce precision for privacy, make that an explicit product/security policy and test its effect on query outcomes.

5. Controlled failure: remove the required geospatial index

prove that proximity operators depend on the index, then restore it
const d=db.getSiblingDB("atlasmart");d.pickup_points_ch20_l5.dropIndex("tenant_open_geo");try { d.pickup_points_ch20_l5.find({location:{$near:{$geometry:{type:"Point",coordinates:[49.8671,40.4093]},$maxDistance:2000}}}).limit(5).toArray();} catch(e){ print("expected missing-geospatial-index failure",e.code,e.codeName); }d.pickup_points_ch20_l5.createIndex({tenantId:1,open:1,location:"2dsphere"},{name:"tenant_open_geo"});printjson(d.pickup_points_ch20_l5.getIndexes());

The repair is not “add every possible geo index.” Restore the one index that matches a justified query contract, then re-measure writes, index bytes, and reads. Extra indexes amplify write cost and cache/storage pressure.

6. Production judgment and regression checklist

Operate a location service with explicit SLOs and abuse limits. Monitor request radius distribution, result counts, p50/p95/p99 latency, timeout/error rates, keys/documents examined, index size, write rate, cache pressure, malformed/rejected coordinates, and geographic outliers. Regression tests should include a known nearest ordering, exact tenant isolation, points just inside/outside service polygons, antimeridian/polar fixtures when relevant, multiple geospatial-index selection, and the no-index failure path.

For sharded collections, the geospatial index cannot itself be the shard key; route/shard by another business key and test scatter/targeting separately. If the application later needs relevance-ranked text plus geography, Chapter 21 will distinguish database geospatial indexes from dedicated MongoDB Search geo operators and their different scoring/geometry semantics.

Bridge. Chapter 21 moves from deterministic database predicates to dedicated Search and Vector Search, where analyzers, relevance, faceting, ANN recall, and search-specific geo behavior introduce a different operational plane.

cleanup only this lesson lab
docker rm -f atlasmart-ch20-l5 2>/dev/null || truedocker volume rm atlasmart-ch20-l5-data 2>/dev/null || true

Check your understanding

  1. Why cap radius and result count at the API boundary?
  2. Why are benchmark latency values not pre-filled in the lesson?
  3. What does coordinate-rounding drift measure?
  4. What happens if $near has no suitable geospatial index?
  5. Why not add many geospatial indexes preemptively?
Review the answers

1. They bound candidate work, response size, and abuse potential.

2. They depend on hardware, cache, storage, client, concurrency, and workload; invented numbers would be misleading.

3. Physical displacement introduced by reducing decimal precision for a specific point.

4. The query fails; proximity operators require a geospatial index.

5. Each index adds write, storage, and cache cost; create indexes for measured query contracts.

Authoritative references

  • Geospatial Queries — GeoJSON versus legacy coordinates, WGS84, operators, and MongoDB 8.2 representation precedence.
  • GeoJSON Objects — supported geometry forms and longitude/latitude ordering.
  • 2dsphere Indexes — spherical indexing, sparse behavior, compound rules, and version 4 in MongoDB 8.3.
  • 2d Indexes — flat Euclidean indexing for legacy coordinate pairs and compound limitations.
  • Geospatial Index Restrictions — covered-query, collation, shard-key, supported-data, and multiple-index restrictions.
  • $geometry — EPSG:4326 default coordinate-reference behavior and GeoJSON geometry syntax.
  • $near — nearest-first query semantics and index requirements.
  • $nearSphere — spherical proximity, distance units, validation, and sorting behavior.
  • $geoWithin — containment predicates and unsorted spatial filtering.
  • $geoIntersects — intersection predicates for GeoJSON geometry.
  • $geoNear — first-stage/index rules, distanceField/key/query options, and distance units.
  • $centerSphere — spherical-cap radius semantics in radians.
  • Explain Results — executionStats interpretation and plan-format caveats.
  • MongoDB 8.3 Release Notes — 2dsphere index version 4 and current 8.3 behavior.
  • mongosh Release Notes — mongosh 2.10.0 baseline.
  • PyMongo Release Notes — PyMongo 4.17 driver baseline.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.