[{"data":1,"prerenderedAt":2929},["ShallowReactive",2],{"page-\u002Fscaling-and-operating-production-python-apis\u002Fmonitoring-and-logging-python-apis\u002F":3,"faq-schema-\u002Fscaling-and-operating-production-python-apis\u002Fmonitoring-and-logging-python-apis\u002F":2908},{"id":4,"title":5,"body":6,"description":2899,"extension":2900,"meta":2901,"navigation":550,"path":2904,"seo":2905,"stem":2906,"__hash__":2907},"content\u002Fscaling-and-operating-production-python-apis\u002Fmonitoring-and-logging-python-apis\u002Findex.md","Monitoring and Logging Python APIs in Production",{"type":7,"value":8,"toc":2884},"minimark",[9,13,28,40,164,167,172,188,194,293,296,298,302,310,360,378,464,472,474,478,498,508,741,747,1072,1086,1152,1174,1176,1180,1183,1595,1598,1601,1785,1803,1805,1809,1812,1819,1822,1915,1918,1921,2123,2144,2146,2150,2153,2159,2170,2172,2176,2184,2405,2411,2413,2417,2514,2516,2520,2575,2577,2581,2584,2619,2625,2722,2741,2743,2747,2750,2768,2770,2774,2786,2795,2805,2814,2824,2826,2830,2835,2861,2866,2880],[10,11,5],"h1",{"id":12},"monitoring-and-logging-python-apis-in-production",[14,15,16,17,21,22,27],"p",{},"An API that runs without instrumentation is an API you cannot operate. When a customer reports \"it's slow\" or your Stripe bill spikes overnight, you need a ",[18,19,20],"code",{},"request_id"," to grep, a latency histogram to query, and a cost-per-request number to defend your margins — not a guess and a shrug. This guide wires structured logging, Prometheus metrics, and the four numbers that actually matter into a FastAPI service, then shows you how to turn those numbers into alerts before a customer turns them into a churn event. Part of the ",[23,24,26],"a",{"href":25},"\u002Fscaling-and-operating-production-python-apis\u002F","Scaling and Operating Production Python APIs"," guide.",[14,29,30,31,34,35,39],{},"The goal here is operational leverage: one consistent log schema you can query, a ",[18,32,33],{},"\u002Fmetrics"," endpoint your dashboard scrapes, and enough cost visibility that you can tie a slow endpoint back to an eroded gross margin on your ",[23,36,38],{"href":37},"\u002Fbuilding-monetizing-api-driven-micro-saas\u002Fdesigning-api-pricing-tiers\u002F","pricing tiers",". Observability is not a nice-to-have you bolt on after launch; for a one- or two-person team it is the difference between debugging in minutes and debugging by customer email. You do not have a support rotation or an SRE team to absorb the cost of flying blind, so the instrumentation has to do that work for you.",[41,42,50,51,50,55,50,59,50,66,50,85,50,94,50,102,50,109,50,114,50,120,50,125,50,129,50,134,50,137,50,143,50,148,50,152,50,156,50,160],"svg",{"viewBox":43,"role":44,"ariaLabelledBy":45,"xmlns":48,"style":49},"0 0 720 240","img",[46,47],"monlog-flow-t","monlog-flow-d","http:\u002F\u002Fwww.w3.org\u002F2000\u002Fsvg","width:100%;height:auto;margin:1.5rem 0;font-family:var(--font-sans);","\n  ",[52,53,54],"title",{"id":46},"Request flow through logging and metrics middleware",[56,57,58],"desc",{"id":47},"A client request passes through middleware that emits a structured JSON log line and records latency metrics, both flowing to a log sink and a metrics store feeding a dashboard.",[60,61],"rect",{"x":62,"y":62,"width":63,"height":64,"fill":65},"0","720","240","var(--c-surface)",[67,68,69,70,50],"defs",{},"\n    ",[71,72,79,80,69],"marker",{"id":73,"viewBox":74,"refX":75,"refY":76,"markerWidth":77,"markerHeight":77,"orient":78},"monlog-flow-arrow","0 0 10 10","9","5","7","auto-start-reverse","\n      ",[81,82],"path",{"d":83,"fill":84},"M0 0 L10 5 L0 10 z","var(--c-text-muted)",[60,86],{"x":87,"y":88,"width":89,"height":90,"rx":91,"fill":65,"stroke":92,"style":93},"20","90","150","60","10","var(--c-blue)","stroke-width:2;",[95,96,101],"text",{"x":97,"y":98,"fill":99,"style":100},"95","125","var(--c-text)","text-anchor:middle;font-size:14;font-family:var(--font-sans);","Client request",[103,104],"line",{"x1":105,"y1":106,"x2":107,"y2":106,"stroke":84,"style":108},"170","120","230","stroke-width:2;marker-end:url(#monlog-flow-arrow);",[60,110],{"x":107,"y":111,"width":112,"height":111,"rx":91,"fill":65,"stroke":113,"style":93},"80","180","var(--c-teal)",[95,115,119],{"x":116,"y":117,"fill":99,"style":118},"320","108","text-anchor:middle;font-size:13;font-family:var(--font-sans);","Middleware",[95,121,124],{"x":116,"y":122,"fill":84,"style":123},"126","text-anchor:middle;font-size:12;font-family:var(--font-sans);","request_id + timer",[95,126,128],{"x":116,"y":127,"fill":84,"style":123},"144","JSON log + histogram",[103,130],{"x1":131,"y1":132,"x2":133,"y2":90,"stroke":84,"style":108},"410","105","490",[103,135],{"x1":131,"y1":136,"x2":133,"y2":112,"stroke":84,"style":108},"135",[60,138],{"x":133,"y":139,"width":140,"height":141,"rx":91,"fill":65,"stroke":142,"style":93},"30","210","58","var(--c-yellow)",[95,144,147],{"x":145,"y":146,"fill":99,"style":118},"595","55","Log sink",[95,149,151],{"x":145,"y":150,"fill":84,"style":123},"73","Loki \u002F CloudWatch",[60,153],{"x":133,"y":154,"width":140,"height":141,"rx":91,"fill":65,"stroke":155,"style":93},"152","var(--c-coral)",[95,157,159],{"x":145,"y":158,"fill":99,"style":118},"177","Metrics store",[95,161,163],{"x":145,"y":162,"fill":84,"style":123},"195","Prometheus + Grafana",[165,166],"hr",{},[168,169,171],"h2",{"id":170},"the-three-signals-logs-metrics-and-traces","The three signals: logs, metrics, and traces",[14,173,174,175,179,180,183,184,187],{},"Before writing any code, get the mental model right, because most builders overspend on one signal and ignore the other two. Observability rests on three distinct signals, and they answer different questions. ",[176,177,178],"strong",{},"Logs"," are discrete events with full context — they answer \"what exactly happened to this one request.\" ",[176,181,182],{},"Metrics"," are cheap numeric aggregates sampled over time — they answer \"what is happening across all requests right now.\" ",[176,185,186],{},"Traces"," stitch a single request's path across services — they answer \"which span inside this slow request was actually slow.\" Reaching for the wrong one wastes money: paging through logs to estimate a p95 is slow and imprecise, and firing a metric per user id will bankrupt your Prometheus server.",[14,189,190,191,193],{},"The reason to run all three is that each one is nearly useless alone. Metrics show you a spike at 14:32 but cannot tell you which customer or payload caused it. Logs let you reconstruct that one request in detail but cannot show you the trend that told you to look. Traces pinpoint the slow database call but only once metrics have told you a route is slow and logs have handed you the offending ",[18,192,20],{},". The craft is matching the signal to the question and keeping each one's cost proportional to its value.",[41,195,50,200,50,203,50,206,50,209,50,213,50,217,50,221,50,226,50,230,50,234,50,238,50,242,50,245,50,248,50,250,50,254,50,258,50,262,50,265,50,268,50,271,50,274,50,278,50,282,50,285,50,291],{"viewBox":196,"role":44,"ariaLabelledBy":197,"xmlns":48,"style":49},"0 0 720 300",[198,199],"monlog-signals-t","monlog-signals-d",[52,201,202],{"id":198},"Comparison of logs, metrics, and traces",[56,204,205],{"id":199},"A matrix comparing the three observability signals by the question each answers, its cardinality tolerance, and its main cost driver.",[60,207],{"x":62,"y":62,"width":63,"height":208,"fill":65},"300",[95,210,212],{"x":89,"y":211,"fill":99,"style":118},"40","Answers",[95,214,216],{"x":215,"y":211,"fill":99,"style":118},"390","Cardinality",[95,218,220],{"x":219,"y":211,"fill":99,"style":118},"600","Cost driver",[60,222],{"x":87,"y":90,"width":223,"height":224,"rx":225,"fill":65,"stroke":92,"style":93},"110","64","8",[95,227,178],{"x":228,"y":229,"fill":99,"style":100},"75","97",[95,231,233],{"x":89,"y":232,"fill":84,"style":123},"88","One request,",[95,235,237],{"x":89,"y":236,"fill":84,"style":123},"104","full detail",[95,239,241],{"x":215,"y":240,"fill":84,"style":123},"96","High is fine",[95,243,244],{"x":219,"y":240,"fill":84,"style":123},"Volume stored",[60,246],{"x":87,"y":247,"width":223,"height":224,"rx":225,"fill":65,"stroke":113,"style":93},"140",[95,249,182],{"x":228,"y":158,"fill":99,"style":100},[95,251,253],{"x":89,"y":252,"fill":84,"style":123},"168","All requests,",[95,255,257],{"x":89,"y":256,"fill":84,"style":123},"184","trends only",[95,259,261],{"x":215,"y":260,"fill":84,"style":123},"176","Keep it low",[95,263,264],{"x":219,"y":260,"fill":84,"style":123},"Series count",[60,266],{"x":87,"y":267,"width":223,"height":224,"rx":225,"fill":65,"stroke":155,"style":93},"220",[95,269,186],{"x":228,"y":270,"fill":99,"style":100},"257",[95,272,233],{"x":89,"y":273,"fill":84,"style":123},"248",[95,275,277],{"x":89,"y":276,"fill":84,"style":123},"264","span by span",[95,279,281],{"x":215,"y":280,"fill":84,"style":123},"256","Sample hard",[95,283,284],{"x":219,"y":280,"fill":84,"style":123},"Export volume",[103,286],{"x1":287,"y1":146,"x2":287,"y2":288,"stroke":289,"style":290},"185","290","var(--c-border)","stroke-width:1;",[103,292],{"x1":133,"y1":146,"x2":133,"y2":288,"stroke":289,"style":290},[14,294,295],{},"For most APIs the right order of investment is logs first, metrics second, traces last. Logs cost nothing to start and pay off on day one when a webhook fails. Metrics are next because they unlock alerting and capacity planning. Traces come last because they earn their keep only once your request fans out across a database, a cache, and one or more upstream APIs — before that, a good log line already tells you where the time went.",[165,297],{},[168,299,301],{"id":300},"prerequisites","Prerequisites",[14,303,304,305,309],{},"This guide assumes a working FastAPI service on Python 3.11+ and that you are comfortable with ",[23,306,308],{"href":307},"\u002Fgetting-started-with-python-apis-for-builders\u002Fmaking-http-requests-with-requests-library\u002F","making HTTP requests"," and async route handlers. Install the two instrumentation libraries:",[311,312,317],"pre",{"className":313,"code":314,"language":315,"meta":316,"style":316},"language-bash shiki shiki-themes github-light github-dark","pip install structlog prometheus-client\n# optional, for the tracing step:\npip install opentelemetry-sdk opentelemetry-exporter-otlp opentelemetry-instrumentation-fastapi\n","bash","",[18,318,319,337,344],{"__ignoreMap":316},[320,321,323,327,331,334],"span",{"class":103,"line":322},1,[320,324,326],{"class":325},"sScJk","pip",[320,328,330],{"class":329},"sZZnC"," install",[320,332,333],{"class":329}," structlog",[320,335,336],{"class":329}," prometheus-client\n",[320,338,340],{"class":103,"line":339},2,[320,341,343],{"class":342},"sJ8bj","# optional, for the tracing step:\n",[320,345,347,349,351,354,357],{"class":103,"line":346},3,[320,348,326],{"class":325},[320,350,330],{"class":329},[320,352,353],{"class":329}," opentelemetry-sdk",[320,355,356],{"class":329}," opentelemetry-exporter-otlp",[320,358,359],{"class":329}," opentelemetry-instrumentation-fastapi\n",[14,361,362,363,369,370,373,374,377],{},"We use ",[23,364,366],{"href":365},"\u002Fscaling-and-operating-production-python-apis\u002Fmonitoring-and-logging-python-apis\u002Fstructured-logging-with-structlog\u002F",[18,367,368],{},"structlog"," for log emission, but every snippet works with the standard library ",[18,371,372],{},"logging"," module configured with a JSON formatter — the ",[23,375,376],{"href":365},"dedicated structlog guide"," covers the trade-off. Configuration is driven entirely by environment variables, never hardcoded, so the same image runs identically in local development, staging, and production with only the environment changing:",[379,380,381,397],"table",{},[382,383,384],"thead",{},[385,386,387,391,394],"tr",{},[388,389,390],"th",{},"Env var",[388,392,393],{},"Purpose",[388,395,396],{},"Example",[398,399,400,416,431,446],"tbody",{},[385,401,402,408,411],{},[403,404,405],"td",{},[18,406,407],{},"LOG_LEVEL",[403,409,410],{},"Minimum log severity emitted",[403,412,413],{},[18,414,415],{},"INFO",[385,417,418,423,426],{},[403,419,420],{},[18,421,422],{},"SERVICE_NAME",[403,424,425],{},"Tag attached to every log and metric",[403,427,428],{},[18,429,430],{},"builder-api",[385,432,433,438,441],{},[403,434,435],{},[18,436,437],{},"OTEL_EXPORTER_OTLP_ENDPOINT",[403,439,440],{},"OpenTelemetry collector endpoint",[403,442,443],{},[18,444,445],{},"http:\u002F\u002Flocalhost:4317",[385,447,448,453,459],{},[403,449,450],{},[18,451,452],{},"METRICS_ENABLED",[403,454,455,456,458],{},"Toggle the ",[18,457,33],{}," endpoint",[403,460,461],{},[18,462,463],{},"true",[14,465,466,467,471],{},"If you deploy in a container, none of this changes — you log to stdout and let the runtime capture the stream, which is exactly the pattern covered in ",[23,468,470],{"href":469},"\u002Fscaling-and-operating-production-python-apis\u002Fcontainerizing-python-apis-with-docker\u002F","containerizing Python APIs with Docker",".",[165,473],{},[168,475,477],{"id":476},"step-1-json-structured-logging-with-a-request_id","Step 1: JSON structured logging with a request_id",[14,479,480,481,484,485,487,488,490,491,494,495,497],{},"Plain-text logs are unqueryable. The moment you have two concurrent requests, interleaved ",[18,482,483],{},"print()"," output becomes useless for tracing a single call — line 3 belongs to request A, line 4 to request B, and you have no way to tell them apart. The fix is one JSON object per line, every line carrying a ",[18,486,20],{}," that ties together the entry, exit, and any errors of a single request. Once every line is structured JSON, your log backend can filter by ",[18,489,20],{},", ",[18,492,493],{},"status",", or ",[18,496,81],{}," the way a database filters rows, and the interleaving stops mattering.",[14,499,500,501,503,504,507],{},"Bind the ",[18,502,20],{}," into a ",[18,505,506],{},"contextvars","-backed context so every log call inside the request automatically inherits it, with no parameter threading:",[311,509,513],{"className":510,"code":511,"language":512,"meta":316,"style":316},"language-python shiki shiki-themes github-light github-dark","# logging_setup.py\nimport logging\nimport os\nimport structlog\n\ndef configure_logging() -> None:\n    level = os.getenv(\"LOG_LEVEL\", \"INFO\").upper()\n    logging.basicConfig(format=\"%(message)s\", level=level)\n    structlog.configure(\n        processors=[\n            structlog.contextvars.merge_contextvars,\n            structlog.processors.add_log_level,\n            structlog.processors.TimeStamper(fmt=\"iso\"),\n            structlog.processors.JSONRenderer(),\n        ],\n        wrapper_class=structlog.make_filtering_bound_logger(\n            logging.getLevelName(level)\n        ),\n        cache_logger_on_first_use=True,\n    )\n\nlog = structlog.get_logger()\n","python",[18,514,515,520,530,537,545,552,571,594,624,630,641,647,653,670,676,682,693,699,705,719,725,730],{"__ignoreMap":316},[320,516,517],{"class":103,"line":322},[320,518,519],{"class":342},"# logging_setup.py\n",[320,521,522,526],{"class":103,"line":339},[320,523,525],{"class":524},"szBVR","import",[320,527,529],{"class":528},"sVt8B"," logging\n",[320,531,532,534],{"class":103,"line":346},[320,533,525],{"class":524},[320,535,536],{"class":528}," os\n",[320,538,540,542],{"class":103,"line":539},4,[320,541,525],{"class":524},[320,543,544],{"class":528}," structlog\n",[320,546,548],{"class":103,"line":547},5,[320,549,551],{"emptyLinePlaceholder":550},true,"\n",[320,553,555,558,561,564,568],{"class":103,"line":554},6,[320,556,557],{"class":524},"def",[320,559,560],{"class":325}," configure_logging",[320,562,563],{"class":528},"() -> ",[320,565,567],{"class":566},"sj4cs","None",[320,569,570],{"class":528},":\n",[320,572,574,577,580,583,586,588,591],{"class":103,"line":573},7,[320,575,576],{"class":528},"    level ",[320,578,579],{"class":524},"=",[320,581,582],{"class":528}," os.getenv(",[320,584,585],{"class":329},"\"LOG_LEVEL\"",[320,587,490],{"class":528},[320,589,590],{"class":329},"\"INFO\"",[320,592,593],{"class":528},").upper()\n",[320,595,597,600,604,606,609,612,614,616,619,621],{"class":103,"line":596},8,[320,598,599],{"class":528},"    logging.basicConfig(",[320,601,603],{"class":602},"s4XuR","format",[320,605,579],{"class":524},[320,607,608],{"class":329},"\"",[320,610,611],{"class":566},"%(message)s",[320,613,608],{"class":329},[320,615,490],{"class":528},[320,617,618],{"class":602},"level",[320,620,579],{"class":524},[320,622,623],{"class":528},"level)\n",[320,625,627],{"class":103,"line":626},9,[320,628,629],{"class":528},"    structlog.configure(\n",[320,631,633,636,638],{"class":103,"line":632},10,[320,634,635],{"class":602},"        processors",[320,637,579],{"class":524},[320,639,640],{"class":528},"[\n",[320,642,644],{"class":103,"line":643},11,[320,645,646],{"class":528},"            structlog.contextvars.merge_contextvars,\n",[320,648,650],{"class":103,"line":649},12,[320,651,652],{"class":528},"            structlog.processors.add_log_level,\n",[320,654,656,659,662,664,667],{"class":103,"line":655},13,[320,657,658],{"class":528},"            structlog.processors.TimeStamper(",[320,660,661],{"class":602},"fmt",[320,663,579],{"class":524},[320,665,666],{"class":329},"\"iso\"",[320,668,669],{"class":528},"),\n",[320,671,673],{"class":103,"line":672},14,[320,674,675],{"class":528},"            structlog.processors.JSONRenderer(),\n",[320,677,679],{"class":103,"line":678},15,[320,680,681],{"class":528},"        ],\n",[320,683,685,688,690],{"class":103,"line":684},16,[320,686,687],{"class":602},"        wrapper_class",[320,689,579],{"class":524},[320,691,692],{"class":528},"structlog.make_filtering_bound_logger(\n",[320,694,696],{"class":103,"line":695},17,[320,697,698],{"class":528},"            logging.getLevelName(level)\n",[320,700,702],{"class":103,"line":701},18,[320,703,704],{"class":528},"        ),\n",[320,706,708,711,713,716],{"class":103,"line":707},19,[320,709,710],{"class":602},"        cache_logger_on_first_use",[320,712,579],{"class":524},[320,714,715],{"class":566},"True",[320,717,718],{"class":528},",\n",[320,720,722],{"class":103,"line":721},20,[320,723,724],{"class":528},"    )\n",[320,726,728],{"class":103,"line":727},21,[320,729,551],{"emptyLinePlaceholder":550},[320,731,733,736,738],{"class":103,"line":732},22,[320,734,735],{"class":528},"log ",[320,737,579],{"class":524},[320,739,740],{"class":528}," structlog.get_logger()\n",[14,742,743,744,746],{},"Now add middleware that generates a ",[18,745,20],{},", binds it, and logs request completion with its status and duration:",[311,748,750],{"className":510,"code":749,"language":512,"meta":316,"style":316},"# middleware.py\nimport os\nimport time\nimport uuid\nimport structlog\nfrom fastapi import Request\n\nlog = structlog.get_logger()\nSERVICE_NAME = os.getenv(\"SERVICE_NAME\", \"builder-api\")\n\nasync def logging_middleware(request: Request, call_next):\n    request_id = request.headers.get(\"X-Request-ID\", str(uuid.uuid4()))\n    structlog.contextvars.bind_contextvars(\n        request_id=request_id,\n        service=SERVICE_NAME,\n        path=request.url.path,\n        method=request.method,\n    )\n    start = time.perf_counter()\n    try:\n        response = await call_next(request)\n    except Exception:\n        log.exception(\"request_failed\")\n        structlog.contextvars.clear_contextvars()\n        raise\n    duration_ms = (time.perf_counter() - start) * 1000\n    log.info(\"request_completed\", status=response.status_code,\n             duration_ms=round(duration_ms, 2))\n    response.headers[\"X-Request-ID\"] = request_id\n    structlog.contextvars.clear_contextvars()\n    return response\n",[18,751,752,757,763,770,777,783,796,800,808,828,832,846,867,872,882,893,903,913,917,927,934,947,957,968,974,980,1003,1021,1041,1057,1063],{"__ignoreMap":316},[320,753,754],{"class":103,"line":322},[320,755,756],{"class":342},"# middleware.py\n",[320,758,759,761],{"class":103,"line":339},[320,760,525],{"class":524},[320,762,536],{"class":528},[320,764,765,767],{"class":103,"line":346},[320,766,525],{"class":524},[320,768,769],{"class":528}," time\n",[320,771,772,774],{"class":103,"line":539},[320,773,525],{"class":524},[320,775,776],{"class":528}," uuid\n",[320,778,779,781],{"class":103,"line":547},[320,780,525],{"class":524},[320,782,544],{"class":528},[320,784,785,788,791,793],{"class":103,"line":554},[320,786,787],{"class":524},"from",[320,789,790],{"class":528}," fastapi ",[320,792,525],{"class":524},[320,794,795],{"class":528}," Request\n",[320,797,798],{"class":103,"line":573},[320,799,551],{"emptyLinePlaceholder":550},[320,801,802,804,806],{"class":103,"line":596},[320,803,735],{"class":528},[320,805,579],{"class":524},[320,807,740],{"class":528},[320,809,810,812,815,817,820,822,825],{"class":103,"line":626},[320,811,422],{"class":566},[320,813,814],{"class":524}," =",[320,816,582],{"class":528},[320,818,819],{"class":329},"\"SERVICE_NAME\"",[320,821,490],{"class":528},[320,823,824],{"class":329},"\"builder-api\"",[320,826,827],{"class":528},")\n",[320,829,830],{"class":103,"line":632},[320,831,551],{"emptyLinePlaceholder":550},[320,833,834,837,840,843],{"class":103,"line":643},[320,835,836],{"class":524},"async",[320,838,839],{"class":524}," def",[320,841,842],{"class":325}," logging_middleware",[320,844,845],{"class":528},"(request: Request, call_next):\n",[320,847,848,851,853,856,859,861,864],{"class":103,"line":649},[320,849,850],{"class":528},"    request_id ",[320,852,579],{"class":524},[320,854,855],{"class":528}," request.headers.get(",[320,857,858],{"class":329},"\"X-Request-ID\"",[320,860,490],{"class":528},[320,862,863],{"class":566},"str",[320,865,866],{"class":528},"(uuid.uuid4()))\n",[320,868,869],{"class":103,"line":655},[320,870,871],{"class":528},"    structlog.contextvars.bind_contextvars(\n",[320,873,874,877,879],{"class":103,"line":672},[320,875,876],{"class":602},"        request_id",[320,878,579],{"class":524},[320,880,881],{"class":528},"request_id,\n",[320,883,884,887,889,891],{"class":103,"line":678},[320,885,886],{"class":602},"        service",[320,888,579],{"class":524},[320,890,422],{"class":566},[320,892,718],{"class":528},[320,894,895,898,900],{"class":103,"line":684},[320,896,897],{"class":602},"        path",[320,899,579],{"class":524},[320,901,902],{"class":528},"request.url.path,\n",[320,904,905,908,910],{"class":103,"line":695},[320,906,907],{"class":602},"        method",[320,909,579],{"class":524},[320,911,912],{"class":528},"request.method,\n",[320,914,915],{"class":103,"line":701},[320,916,724],{"class":528},[320,918,919,922,924],{"class":103,"line":707},[320,920,921],{"class":528},"    start ",[320,923,579],{"class":524},[320,925,926],{"class":528}," time.perf_counter()\n",[320,928,929,932],{"class":103,"line":721},[320,930,931],{"class":524},"    try",[320,933,570],{"class":528},[320,935,936,939,941,944],{"class":103,"line":727},[320,937,938],{"class":528},"        response ",[320,940,579],{"class":524},[320,942,943],{"class":524}," await",[320,945,946],{"class":528}," call_next(request)\n",[320,948,949,952,955],{"class":103,"line":732},[320,950,951],{"class":524},"    except",[320,953,954],{"class":566}," Exception",[320,956,570],{"class":528},[320,958,960,963,966],{"class":103,"line":959},23,[320,961,962],{"class":528},"        log.exception(",[320,964,965],{"class":329},"\"request_failed\"",[320,967,827],{"class":528},[320,969,971],{"class":103,"line":970},24,[320,972,973],{"class":528},"        structlog.contextvars.clear_contextvars()\n",[320,975,977],{"class":103,"line":976},25,[320,978,979],{"class":524},"        raise\n",[320,981,983,986,988,991,994,997,1000],{"class":103,"line":982},26,[320,984,985],{"class":528},"    duration_ms ",[320,987,579],{"class":524},[320,989,990],{"class":528}," (time.perf_counter() ",[320,992,993],{"class":524},"-",[320,995,996],{"class":528}," start) ",[320,998,999],{"class":524},"*",[320,1001,1002],{"class":566}," 1000\n",[320,1004,1006,1009,1012,1014,1016,1018],{"class":103,"line":1005},27,[320,1007,1008],{"class":528},"    log.info(",[320,1010,1011],{"class":329},"\"request_completed\"",[320,1013,490],{"class":528},[320,1015,493],{"class":602},[320,1017,579],{"class":524},[320,1019,1020],{"class":528},"response.status_code,\n",[320,1022,1024,1027,1029,1032,1035,1038],{"class":103,"line":1023},28,[320,1025,1026],{"class":602},"             duration_ms",[320,1028,579],{"class":524},[320,1030,1031],{"class":566},"round",[320,1033,1034],{"class":528},"(duration_ms, ",[320,1036,1037],{"class":566},"2",[320,1039,1040],{"class":528},"))\n",[320,1042,1044,1047,1049,1052,1054],{"class":103,"line":1043},29,[320,1045,1046],{"class":528},"    response.headers[",[320,1048,858],{"class":329},[320,1050,1051],{"class":528},"] ",[320,1053,579],{"class":524},[320,1055,1056],{"class":528}," request_id\n",[320,1058,1060],{"class":103,"line":1059},30,[320,1061,1062],{"class":528},"    structlog.contextvars.clear_contextvars()\n",[320,1064,1066,1069],{"class":103,"line":1065},31,[320,1067,1068],{"class":524},"    return",[320,1070,1071],{"class":528}," response\n",[14,1073,1074,1075,1078,1079,1082,1083,1085],{},"Two details in that middleware are load-bearing. First, it honours an inbound ",[18,1076,1077],{},"X-Request-ID"," header before generating its own, so when a request arrives from an upstream gateway or a webhook sender that already assigned an id, the trail stays continuous across service boundaries. Second, it clears the context on both the success and the exception path — skip the ",[18,1080,1081],{},"clear_contextvars()"," in the error branch and a leaked ",[18,1084,20],{}," will bleed into the next request that happens to reuse the same worker, poisoning your logs with the wrong id.",[41,1087,50,1092,50,1095,50,1098,50,1100,50,1103,50,1106,50,1109,50,1113,50,1116,50,1119,50,1122,50,1125,50,1128,50,1130,50,1132,50,1135,50,1138,50,1141,50,1144,50,1147],{"viewBox":1088,"role":44,"ariaLabelledBy":1089,"xmlns":48,"style":49},"0 0 720 210",[1090,1091],"monlog-life-t","monlog-life-d",[52,1093,1094],{"id":1090},"Lifecycle of a request_id inside the logging middleware",[56,1096,1097],{"id":1091},"A timeline showing a request_id bound at request entry, inherited by every log line during handling, attached to the completion log, then cleared at exit.",[60,1099],{"x":62,"y":62,"width":63,"height":140,"fill":65},[103,1101],{"x1":211,"y1":106,"x2":1102,"y2":106,"stroke":289,"style":93},"680",[1104,1105],"circle",{"cx":88,"cy":106,"r":75,"fill":65,"stroke":92,"style":93},[95,1107,1108],{"x":88,"y":97,"fill":99,"style":123},"bind",[95,1110,1112],{"x":88,"y":89,"fill":84,"style":1111},"text-anchor:middle;font-size:11;font-family:var(--font-sans);","id created",[1104,1114],{"cx":1115,"cy":106,"r":75,"fill":65,"stroke":113,"style":93},"270",[95,1117,1118],{"x":1115,"y":97,"fill":99,"style":123},"handler logs",[95,1120,1121],{"x":1115,"y":89,"fill":84,"style":1111},"id inherited",[1104,1123],{"cx":1124,"cy":106,"r":75,"fill":65,"stroke":113,"style":93},"450",[95,1126,1127],{"x":1124,"y":97,"fill":99,"style":123},"more logs",[95,1129,1121],{"x":1124,"y":89,"fill":84,"style":1111},[1104,1131],{"cx":219,"cy":106,"r":75,"fill":65,"stroke":142,"style":93},[95,1133,1134],{"x":219,"y":97,"fill":99,"style":123},"completed",[95,1136,1137],{"x":219,"y":89,"fill":84,"style":1111},"+status\u002Fms",[1104,1139],{"cx":1140,"cy":106,"r":75,"fill":65,"stroke":155,"style":93},"660",[95,1142,1143],{"x":1140,"y":287,"fill":84,"style":1111},"clear",[60,1145],{"x":87,"y":139,"width":1102,"height":139,"rx":1146,"fill":65,"stroke":289,"style":290},"6",[95,1148,1151],{"x":1149,"y":1150,"fill":84,"style":123},"360","50","request_id present in contextvars for the whole request, then removed",[14,1153,1154,1155,1157,1158,1160,1161,1164,1165,1167,1168,1170,1171,471],{},"Returning the ",[18,1156,1077],{}," header lets a customer quote the exact ID from a failed call, turning a vague bug report into a one-line grep. One caveat with ",[18,1159,506],{}," and async: if you spawn a background coroutine with ",[18,1162,1163],{},"asyncio.create_task"," inside a request, it captures the context at creation time, so the ",[18,1166,20],{}," follows it — but a task you fire and forget after the response has returned will log against a context you have already cleared, so bind a fresh id for genuinely detached work. For the full processor pipeline and why ",[18,1169,506],{}," beats passing loggers around, see ",[23,1172,1173],{"href":365},"structured logging with structlog",[165,1175],{},[168,1177,1179],{"id":1178},"step-2-a-prometheus-metrics-endpoint-with-a-latency-histogram","Step 2: A Prometheus \u002Fmetrics endpoint with a latency histogram",[14,1181,1182],{},"Logs answer \"what happened to this request\"; metrics answer \"what is happening across all requests\". A histogram of request latency is the single most valuable metric you can expose — it powers p95\u002Fp99, alerting, and capacity planning, all from one instrument that costs a handful of counter increments per request.",[311,1184,1186],{"className":510,"code":1185,"language":512,"meta":316,"style":316},"# metrics.py\nimport os\nimport time\nfrom prometheus_client import Counter, Histogram, generate_latest, CONTENT_TYPE_LATEST\nfrom fastapi import Request, Response\n\nREQUEST_LATENCY = Histogram(\n    \"http_request_duration_seconds\",\n    \"Request latency in seconds\",\n    labelnames=(\"method\", \"route\", \"status\"),\n    buckets=(0.01, 0.025, 0.05, 0.1, 0.25, 0.5, 1.0, 2.5, 5.0),\n)\nREQUEST_COUNT = Counter(\n    \"http_requests_total\",\n    \"Total requests\",\n    labelnames=(\"method\", \"route\", \"status\"),\n)\n\nasync def metrics_middleware(request: Request, call_next):\n    start = time.perf_counter()\n    response = await call_next(request)\n    # Use the route template, NOT request.url.path, to avoid high cardinality.\n    route = request.scope.get(\"route\")\n    route_label = route.path if route else \"unmatched\"\n    labels = (request.method, route_label, str(response.status_code))\n    REQUEST_LATENCY.labels(*labels).observe(time.perf_counter() - start)\n    REQUEST_COUNT.labels(*labels).inc()\n    return response\n\nasync def metrics_endpoint(_: Request) -> Response:\n    if os.getenv(\"METRICS_ENABLED\", \"true\").lower() != \"true\":\n        return Response(status_code=404)\n    return Response(generate_latest(), media_type=CONTENT_TYPE_LATEST)\n",[18,1187,1188,1193,1199,1205,1220,1231,1235,1245,1252,1259,1284,1338,1342,1352,1359,1366,1386,1390,1394,1405,1413,1424,1429,1443,1465,1480,1498,1510,1516,1520,1532,1558,1577],{"__ignoreMap":316},[320,1189,1190],{"class":103,"line":322},[320,1191,1192],{"class":342},"# metrics.py\n",[320,1194,1195,1197],{"class":103,"line":339},[320,1196,525],{"class":524},[320,1198,536],{"class":528},[320,1200,1201,1203],{"class":103,"line":346},[320,1202,525],{"class":524},[320,1204,769],{"class":528},[320,1206,1207,1209,1212,1214,1217],{"class":103,"line":539},[320,1208,787],{"class":524},[320,1210,1211],{"class":528}," prometheus_client ",[320,1213,525],{"class":524},[320,1215,1216],{"class":528}," Counter, Histogram, generate_latest, ",[320,1218,1219],{"class":566},"CONTENT_TYPE_LATEST\n",[320,1221,1222,1224,1226,1228],{"class":103,"line":547},[320,1223,787],{"class":524},[320,1225,790],{"class":528},[320,1227,525],{"class":524},[320,1229,1230],{"class":528}," Request, Response\n",[320,1232,1233],{"class":103,"line":554},[320,1234,551],{"emptyLinePlaceholder":550},[320,1236,1237,1240,1242],{"class":103,"line":573},[320,1238,1239],{"class":566},"REQUEST_LATENCY",[320,1241,814],{"class":524},[320,1243,1244],{"class":528}," Histogram(\n",[320,1246,1247,1250],{"class":103,"line":596},[320,1248,1249],{"class":329},"    \"http_request_duration_seconds\"",[320,1251,718],{"class":528},[320,1253,1254,1257],{"class":103,"line":626},[320,1255,1256],{"class":329},"    \"Request latency in seconds\"",[320,1258,718],{"class":528},[320,1260,1261,1264,1266,1269,1272,1274,1277,1279,1282],{"class":103,"line":632},[320,1262,1263],{"class":602},"    labelnames",[320,1265,579],{"class":524},[320,1267,1268],{"class":528},"(",[320,1270,1271],{"class":329},"\"method\"",[320,1273,490],{"class":528},[320,1275,1276],{"class":329},"\"route\"",[320,1278,490],{"class":528},[320,1280,1281],{"class":329},"\"status\"",[320,1283,669],{"class":528},[320,1285,1286,1289,1291,1293,1296,1298,1301,1303,1306,1308,1311,1313,1316,1318,1321,1323,1326,1328,1331,1333,1336],{"class":103,"line":643},[320,1287,1288],{"class":602},"    buckets",[320,1290,579],{"class":524},[320,1292,1268],{"class":528},[320,1294,1295],{"class":566},"0.01",[320,1297,490],{"class":528},[320,1299,1300],{"class":566},"0.025",[320,1302,490],{"class":528},[320,1304,1305],{"class":566},"0.05",[320,1307,490],{"class":528},[320,1309,1310],{"class":566},"0.1",[320,1312,490],{"class":528},[320,1314,1315],{"class":566},"0.25",[320,1317,490],{"class":528},[320,1319,1320],{"class":566},"0.5",[320,1322,490],{"class":528},[320,1324,1325],{"class":566},"1.0",[320,1327,490],{"class":528},[320,1329,1330],{"class":566},"2.5",[320,1332,490],{"class":528},[320,1334,1335],{"class":566},"5.0",[320,1337,669],{"class":528},[320,1339,1340],{"class":103,"line":649},[320,1341,827],{"class":528},[320,1343,1344,1347,1349],{"class":103,"line":655},[320,1345,1346],{"class":566},"REQUEST_COUNT",[320,1348,814],{"class":524},[320,1350,1351],{"class":528}," Counter(\n",[320,1353,1354,1357],{"class":103,"line":672},[320,1355,1356],{"class":329},"    \"http_requests_total\"",[320,1358,718],{"class":528},[320,1360,1361,1364],{"class":103,"line":678},[320,1362,1363],{"class":329},"    \"Total requests\"",[320,1365,718],{"class":528},[320,1367,1368,1370,1372,1374,1376,1378,1380,1382,1384],{"class":103,"line":684},[320,1369,1263],{"class":602},[320,1371,579],{"class":524},[320,1373,1268],{"class":528},[320,1375,1271],{"class":329},[320,1377,490],{"class":528},[320,1379,1276],{"class":329},[320,1381,490],{"class":528},[320,1383,1281],{"class":329},[320,1385,669],{"class":528},[320,1387,1388],{"class":103,"line":695},[320,1389,827],{"class":528},[320,1391,1392],{"class":103,"line":701},[320,1393,551],{"emptyLinePlaceholder":550},[320,1395,1396,1398,1400,1403],{"class":103,"line":707},[320,1397,836],{"class":524},[320,1399,839],{"class":524},[320,1401,1402],{"class":325}," metrics_middleware",[320,1404,845],{"class":528},[320,1406,1407,1409,1411],{"class":103,"line":721},[320,1408,921],{"class":528},[320,1410,579],{"class":524},[320,1412,926],{"class":528},[320,1414,1415,1418,1420,1422],{"class":103,"line":727},[320,1416,1417],{"class":528},"    response ",[320,1419,579],{"class":524},[320,1421,943],{"class":524},[320,1423,946],{"class":528},[320,1425,1426],{"class":103,"line":732},[320,1427,1428],{"class":342},"    # Use the route template, NOT request.url.path, to avoid high cardinality.\n",[320,1430,1431,1434,1436,1439,1441],{"class":103,"line":959},[320,1432,1433],{"class":528},"    route ",[320,1435,579],{"class":524},[320,1437,1438],{"class":528}," request.scope.get(",[320,1440,1276],{"class":329},[320,1442,827],{"class":528},[320,1444,1445,1448,1450,1453,1456,1459,1462],{"class":103,"line":970},[320,1446,1447],{"class":528},"    route_label ",[320,1449,579],{"class":524},[320,1451,1452],{"class":528}," route.path ",[320,1454,1455],{"class":524},"if",[320,1457,1458],{"class":528}," route ",[320,1460,1461],{"class":524},"else",[320,1463,1464],{"class":329}," \"unmatched\"\n",[320,1466,1467,1470,1472,1475,1477],{"class":103,"line":976},[320,1468,1469],{"class":528},"    labels ",[320,1471,579],{"class":524},[320,1473,1474],{"class":528}," (request.method, route_label, ",[320,1476,863],{"class":566},[320,1478,1479],{"class":528},"(response.status_code))\n",[320,1481,1482,1485,1488,1490,1493,1495],{"class":103,"line":982},[320,1483,1484],{"class":566},"    REQUEST_LATENCY",[320,1486,1487],{"class":528},".labels(",[320,1489,999],{"class":524},[320,1491,1492],{"class":528},"labels).observe(time.perf_counter() ",[320,1494,993],{"class":524},[320,1496,1497],{"class":528}," start)\n",[320,1499,1500,1503,1505,1507],{"class":103,"line":1005},[320,1501,1502],{"class":566},"    REQUEST_COUNT",[320,1504,1487],{"class":528},[320,1506,999],{"class":524},[320,1508,1509],{"class":528},"labels).inc()\n",[320,1511,1512,1514],{"class":103,"line":1023},[320,1513,1068],{"class":524},[320,1515,1071],{"class":528},[320,1517,1518],{"class":103,"line":1043},[320,1519,551],{"emptyLinePlaceholder":550},[320,1521,1522,1524,1526,1529],{"class":103,"line":1059},[320,1523,836],{"class":524},[320,1525,839],{"class":524},[320,1527,1528],{"class":325}," metrics_endpoint",[320,1530,1531],{"class":528},"(_: Request) -> Response:\n",[320,1533,1534,1537,1539,1542,1544,1547,1550,1553,1556],{"class":103,"line":1065},[320,1535,1536],{"class":524},"    if",[320,1538,582],{"class":528},[320,1540,1541],{"class":329},"\"METRICS_ENABLED\"",[320,1543,490],{"class":528},[320,1545,1546],{"class":329},"\"true\"",[320,1548,1549],{"class":528},").lower() ",[320,1551,1552],{"class":524},"!=",[320,1554,1555],{"class":329}," \"true\"",[320,1557,570],{"class":528},[320,1559,1561,1564,1567,1570,1572,1575],{"class":103,"line":1560},32,[320,1562,1563],{"class":524},"        return",[320,1565,1566],{"class":528}," Response(",[320,1568,1569],{"class":602},"status_code",[320,1571,579],{"class":524},[320,1573,1574],{"class":566},"404",[320,1576,827],{"class":528},[320,1578,1580,1582,1585,1588,1590,1593],{"class":103,"line":1579},33,[320,1581,1068],{"class":524},[320,1583,1584],{"class":528}," Response(generate_latest(), ",[320,1586,1587],{"class":602},"media_type",[320,1589,579],{"class":524},[320,1591,1592],{"class":566},"CONTENT_TYPE_LATEST",[320,1594,827],{"class":528},[14,1596,1597],{},"The bucket boundaries are a design decision, not a default to accept blindly. Prometheus histograms compute quantiles by interpolating within the bucket the quantile falls into, so a bucket that is too wide makes your p95 imprecise. The buckets above are tuned for a typical JSON API where most responses land between 25 and 250 milliseconds; if your API is a thin proxy that returns in single-digit milliseconds, add finer buckets down low, and if it does heavy work, extend the top end. Pick boundaries that bracket your real target latency so the percentile lands on a tight bucket rather than a wide one.",[14,1599,1600],{},"Wire everything into the app. Order matters: register the metrics middleware before the logging middleware so the timer brackets the whole stack.",[311,1602,1604],{"className":510,"code":1603,"language":512,"meta":316,"style":316},"# main.py\nimport os\nfrom fastapi import FastAPI\nfrom logging_setup import configure_logging\nfrom middleware import logging_middleware\nfrom metrics import metrics_middleware, metrics_endpoint\n\nconfigure_logging()\napp = FastAPI(title=os.getenv(\"SERVICE_NAME\", \"builder-api\"))\napp.middleware(\"http\")(logging_middleware)\napp.middleware(\"http\")(metrics_middleware)\napp.add_route(\"\u002Fmetrics\", metrics_endpoint, include_in_schema=False)\n\n@app.get(\"\u002Fwork\")\nasync def work():\n    return {\"ok\": True}\n",[18,1605,1606,1611,1617,1628,1640,1652,1664,1668,1673,1698,1709,1718,1739,1743,1755,1767],{"__ignoreMap":316},[320,1607,1608],{"class":103,"line":322},[320,1609,1610],{"class":342},"# main.py\n",[320,1612,1613,1615],{"class":103,"line":339},[320,1614,525],{"class":524},[320,1616,536],{"class":528},[320,1618,1619,1621,1623,1625],{"class":103,"line":346},[320,1620,787],{"class":524},[320,1622,790],{"class":528},[320,1624,525],{"class":524},[320,1626,1627],{"class":528}," FastAPI\n",[320,1629,1630,1632,1635,1637],{"class":103,"line":539},[320,1631,787],{"class":524},[320,1633,1634],{"class":528}," logging_setup ",[320,1636,525],{"class":524},[320,1638,1639],{"class":528}," configure_logging\n",[320,1641,1642,1644,1647,1649],{"class":103,"line":547},[320,1643,787],{"class":524},[320,1645,1646],{"class":528}," middleware ",[320,1648,525],{"class":524},[320,1650,1651],{"class":528}," logging_middleware\n",[320,1653,1654,1656,1659,1661],{"class":103,"line":554},[320,1655,787],{"class":524},[320,1657,1658],{"class":528}," metrics ",[320,1660,525],{"class":524},[320,1662,1663],{"class":528}," metrics_middleware, metrics_endpoint\n",[320,1665,1666],{"class":103,"line":573},[320,1667,551],{"emptyLinePlaceholder":550},[320,1669,1670],{"class":103,"line":596},[320,1671,1672],{"class":528},"configure_logging()\n",[320,1674,1675,1678,1680,1683,1685,1687,1690,1692,1694,1696],{"class":103,"line":626},[320,1676,1677],{"class":528},"app ",[320,1679,579],{"class":524},[320,1681,1682],{"class":528}," FastAPI(",[320,1684,52],{"class":602},[320,1686,579],{"class":524},[320,1688,1689],{"class":528},"os.getenv(",[320,1691,819],{"class":329},[320,1693,490],{"class":528},[320,1695,824],{"class":329},[320,1697,1040],{"class":528},[320,1699,1700,1703,1706],{"class":103,"line":632},[320,1701,1702],{"class":528},"app.middleware(",[320,1704,1705],{"class":329},"\"http\"",[320,1707,1708],{"class":528},")(logging_middleware)\n",[320,1710,1711,1713,1715],{"class":103,"line":643},[320,1712,1702],{"class":528},[320,1714,1705],{"class":329},[320,1716,1717],{"class":528},")(metrics_middleware)\n",[320,1719,1720,1723,1726,1729,1732,1734,1737],{"class":103,"line":649},[320,1721,1722],{"class":528},"app.add_route(",[320,1724,1725],{"class":329},"\"\u002Fmetrics\"",[320,1727,1728],{"class":528},", metrics_endpoint, ",[320,1730,1731],{"class":602},"include_in_schema",[320,1733,579],{"class":524},[320,1735,1736],{"class":566},"False",[320,1738,827],{"class":528},[320,1740,1741],{"class":103,"line":655},[320,1742,551],{"emptyLinePlaceholder":550},[320,1744,1745,1748,1750,1753],{"class":103,"line":672},[320,1746,1747],{"class":325},"@app.get",[320,1749,1268],{"class":528},[320,1751,1752],{"class":329},"\"\u002Fwork\"",[320,1754,827],{"class":528},[320,1756,1757,1759,1761,1764],{"class":103,"line":678},[320,1758,836],{"class":524},[320,1760,839],{"class":524},[320,1762,1763],{"class":325}," work",[320,1765,1766],{"class":528},"():\n",[320,1768,1769,1771,1774,1777,1780,1782],{"class":103,"line":684},[320,1770,1068],{"class":524},[320,1772,1773],{"class":528}," {",[320,1775,1776],{"class":329},"\"ok\"",[320,1778,1779],{"class":528},": ",[320,1781,715],{"class":566},[320,1783,1784],{"class":528},"}\n",[14,1786,1787,1788,1791,1792,1795,1796,490,1799,1802],{},"The ",[18,1789,1790],{},"route.path"," template (",[18,1793,1794],{},"\u002Fitems\u002F{id}",") keeps the label count bounded. Using the raw path (",[18,1797,1798],{},"\u002Fitems\u002F42",[18,1800,1801],{},"\u002Fitems\u002F43",", ...) would create one time series per id and melt your Prometheus server — the most common self-inflicted observability outage. A single high-cardinality label can turn a few dozen series into millions in an afternoon, and because Prometheus holds active series in memory, the failure mode is an out-of-memory kill of the whole monitoring stack, not a graceful slowdown. Treat every label value as something you would be comfortable seeing in a dropdown filter.",[165,1804],{},[168,1806,1808],{"id":1807},"step-3-track-p95p99-error-rate-and-cost-per-request","Step 3: Track p95\u002Fp99, error rate, and cost-per-request",[14,1810,1811],{},"The histogram from Step 2 already lets your monitoring backend compute percentiles with a PromQL query — no extra code:",[311,1813,1817],{"className":1814,"code":1816,"language":95,"meta":316},[1815],"language-text","# p95 latency per route over 5 minutes\nhistogram_quantile(0.95, sum(rate(http_request_duration_seconds_bucket[5m])) by (le, route))\n\n# error rate: share of 5xx responses\nsum(rate(http_requests_total{status=~\"5..\"}[5m]))\n  \u002F sum(rate(http_requests_total[5m]))\n",[18,1818,1816],{"__ignoreMap":316},[14,1820,1821],{},"Report p95 and p99, never the average, and this is where most builders go wrong. An average latency hides your worst experiences: if 99 requests return in 40 ms and one takes 4 seconds, the mean is a comfortable-looking 80 ms while a real customer just waited four seconds and may have timed out. The p99 tells you the truth about your tail — the slice of traffic that generates support tickets and churn — and the gap between p50 and p99 tells you how consistent your service is. A p50 of 40 ms with a p99 of 3 s is a service with a serious tail problem hiding behind a healthy median.",[41,1823,50,1827,50,1830,50,1833,50,1835,50,1838,50,1840,50,1843,50,1845,50,1848,50,1852,50,1856,50,1859,50,1861,50,1865,50,1869,50,1873,50,1877,50,1881,50,1885,50,1889,50,1894,50,1898,50,1901,50,1904,50,1907,50,1910],{"viewBox":196,"role":44,"ariaLabelledBy":1824,"xmlns":48,"style":49},[1825,1826],"monlog-hist-t","monlog-hist-d",[52,1828,1829],{"id":1825},"Latency distribution with p95 and p99 markers",[56,1831,1832],{"id":1826},"A bar chart of request counts across latency buckets, showing most traffic under 100 milliseconds with p95 and p99 markers falling in the slow tail.",[60,1834],{"x":62,"y":62,"width":63,"height":208,"fill":65},[103,1836],{"x1":90,"y1":107,"x2":1837,"y2":107,"stroke":289,"style":290},"700",[103,1839],{"x1":90,"y1":211,"x2":90,"y2":107,"stroke":289,"style":290},[95,1841,62],{"x":139,"y":1842,"fill":84,"style":1111},"234",[95,1844,131],{"x":139,"y":224,"fill":84,"style":1111},[60,1846],{"x":1847,"y":112,"width":90,"height":1150,"fill":113},"70",[95,1849,1851],{"x":1850,"y":273,"fill":84,"style":1111},"100","\u003C25ms",[60,1853],{"x":89,"y":1854,"width":90,"height":1855,"fill":113},"89","141",[95,1857,1858],{"x":112,"y":273,"fill":84,"style":1111},"25-50",[60,1860],{"x":107,"y":90,"width":90,"height":105,"fill":92},[95,1862,1864],{"x":1863,"y":273,"fill":84,"style":1111},"260","50-100",[60,1866],{"x":1867,"y":1868,"width":90,"height":228,"fill":113},"310","155",[95,1870,1872],{"x":1871,"y":273,"fill":84,"style":1111},"340","100-250",[60,1874],{"x":215,"y":1875,"width":90,"height":1876,"fill":113},"205","25",[95,1878,1880],{"x":1879,"y":273,"fill":84,"style":1111},"420","250-500",[60,1882],{"x":1883,"y":1884,"width":90,"height":77,"fill":155},"470","223",[95,1886,1888],{"x":1887,"y":273,"fill":84,"style":1111},"500","0.5-1s",[60,1890],{"x":1891,"y":1892,"width":90,"height":1893,"fill":155},"550","226","4",[95,1895,1897],{"x":1896,"y":273,"fill":84,"style":1111},"580",">1s",[103,1899],{"x1":1124,"y1":211,"x2":1124,"y2":107,"stroke":142,"style":1900},"stroke-width:2;stroke-dasharray:5 4;",[95,1902,1903],{"x":1124,"y":139,"fill":99,"style":123},"p95",[103,1905],{"x1":1906,"y1":211,"x2":1906,"y2":107,"stroke":155,"style":1900},"530",[95,1908,1909],{"x":1906,"y":139,"fill":99,"style":123},"p99",[95,1911,1914],{"x":1912,"y":1913,"fill":84,"style":123},"380","282","Request latency bucket (n = 1,132 requests)",[14,1916,1917],{},"The chart above makes the point concrete: the bulk of the 1,132 requests finish under 100 ms, but the p95 sits in the 250–500 ms bucket and the p99 spills into the 500 ms–1 s range. The average would round to roughly 90 ms and quietly ignore the tens of requests that took ten times longer. Alert on the p99, not the mean.",[14,1919,1920],{},"Cost-per-request is the metric that turns observability into a margin defense, and it has to be recorded explicitly because only you know your upstream fees. Track the spend each request incurs — model tokens, third-party API calls, compute — as a counter:",[311,1922,1924],{"className":510,"code":1923,"language":512,"meta":316,"style":316},"# cost.py\nfrom prometheus_client import Counter\n\nREQUEST_COST_USD = Counter(\n    \"request_cost_usd_total\",\n    \"Cumulative request cost in USD\",\n    labelnames=(\"route\", \"tier\"),\n)\n\ndef record_cost(route: str, tier: str, *, upstream_calls: int,\n                tokens: int) -> None:\n    # Wire your real unit economics in here.\n    upstream_fee = upstream_calls * 0.002      # $0.002 per upstream call\n    token_fee = tokens \u002F 1000 * 0.0006         # $0.60 per 1M tokens\n    REQUEST_COST_USD.labels(route=route, tier=tier).inc(\n        upstream_fee + token_fee\n    )\n",[18,1925,1926,1931,1942,1946,1955,1962,1969,1986,1990,1994,2023,2037,2042,2060,2085,2108,2119],{"__ignoreMap":316},[320,1927,1928],{"class":103,"line":322},[320,1929,1930],{"class":342},"# cost.py\n",[320,1932,1933,1935,1937,1939],{"class":103,"line":339},[320,1934,787],{"class":524},[320,1936,1211],{"class":528},[320,1938,525],{"class":524},[320,1940,1941],{"class":528}," Counter\n",[320,1943,1944],{"class":103,"line":346},[320,1945,551],{"emptyLinePlaceholder":550},[320,1947,1948,1951,1953],{"class":103,"line":539},[320,1949,1950],{"class":566},"REQUEST_COST_USD",[320,1952,814],{"class":524},[320,1954,1351],{"class":528},[320,1956,1957,1960],{"class":103,"line":547},[320,1958,1959],{"class":329},"    \"request_cost_usd_total\"",[320,1961,718],{"class":528},[320,1963,1964,1967],{"class":103,"line":554},[320,1965,1966],{"class":329},"    \"Cumulative request cost in USD\"",[320,1968,718],{"class":528},[320,1970,1971,1973,1975,1977,1979,1981,1984],{"class":103,"line":573},[320,1972,1263],{"class":602},[320,1974,579],{"class":524},[320,1976,1268],{"class":528},[320,1978,1276],{"class":329},[320,1980,490],{"class":528},[320,1982,1983],{"class":329},"\"tier\"",[320,1985,669],{"class":528},[320,1987,1988],{"class":103,"line":596},[320,1989,827],{"class":528},[320,1991,1992],{"class":103,"line":626},[320,1993,551],{"emptyLinePlaceholder":550},[320,1995,1996,1998,2001,2004,2006,2009,2011,2013,2015,2018,2021],{"class":103,"line":632},[320,1997,557],{"class":524},[320,1999,2000],{"class":325}," record_cost",[320,2002,2003],{"class":528},"(route: ",[320,2005,863],{"class":566},[320,2007,2008],{"class":528},", tier: ",[320,2010,863],{"class":566},[320,2012,490],{"class":528},[320,2014,999],{"class":524},[320,2016,2017],{"class":528},", upstream_calls: ",[320,2019,2020],{"class":566},"int",[320,2022,718],{"class":528},[320,2024,2025,2028,2030,2033,2035],{"class":103,"line":643},[320,2026,2027],{"class":528},"                tokens: ",[320,2029,2020],{"class":566},[320,2031,2032],{"class":528},") -> ",[320,2034,567],{"class":566},[320,2036,570],{"class":528},[320,2038,2039],{"class":103,"line":649},[320,2040,2041],{"class":342},"    # Wire your real unit economics in here.\n",[320,2043,2044,2047,2049,2052,2054,2057],{"class":103,"line":655},[320,2045,2046],{"class":528},"    upstream_fee ",[320,2048,579],{"class":524},[320,2050,2051],{"class":528}," upstream_calls ",[320,2053,999],{"class":524},[320,2055,2056],{"class":566}," 0.002",[320,2058,2059],{"class":342},"      # $0.002 per upstream call\n",[320,2061,2062,2065,2067,2070,2073,2076,2079,2082],{"class":103,"line":672},[320,2063,2064],{"class":528},"    token_fee ",[320,2066,579],{"class":524},[320,2068,2069],{"class":528}," tokens ",[320,2071,2072],{"class":524},"\u002F",[320,2074,2075],{"class":566}," 1000",[320,2077,2078],{"class":524}," *",[320,2080,2081],{"class":566}," 0.0006",[320,2083,2084],{"class":342},"         # $0.60 per 1M tokens\n",[320,2086,2087,2090,2092,2095,2097,2100,2103,2105],{"class":103,"line":678},[320,2088,2089],{"class":566},"    REQUEST_COST_USD",[320,2091,1487],{"class":528},[320,2093,2094],{"class":602},"route",[320,2096,579],{"class":524},[320,2098,2099],{"class":528},"route, ",[320,2101,2102],{"class":602},"tier",[320,2104,579],{"class":524},[320,2106,2107],{"class":528},"tier).inc(\n",[320,2109,2110,2113,2116],{"class":103,"line":684},[320,2111,2112],{"class":528},"        upstream_fee ",[320,2114,2115],{"class":524},"+",[320,2117,2118],{"class":528}," token_fee\n",[320,2120,2121],{"class":103,"line":695},[320,2122,724],{"class":528},[14,2124,2125,2126,2129,2130,2134,2135,2139,2140,471],{},"Dividing cumulative cost by request count per tier gives average cost-per-request. When that number on your free tier creeps toward your paid tier's price, you have a margin problem before the invoice arrives — the same logic behind ",[23,2127,2128],{"href":37},"usage-based pricing",". The same counter is the raw signal you would persist to a durable store for customer-facing reporting, which is exactly the pattern in ",[23,2131,2133],{"href":2132},"\u002Fbuilding-monetizing-api-driven-micro-saas\u002Ftracking-api-usage-and-analytics\u002Flogging-api-usage-events-to-postgres\u002F","logging API usage events to Postgres",". Watch this number alongside your outbound ",[23,2136,2138],{"href":2137},"\u002Fgetting-started-with-python-apis-for-builders\u002Fmaking-http-requests-with-requests-library\u002Fbest-practices-for-api-rate-limiting\u002F","rate-limiting strategy",", since retries against an upstream multiply both latency and cost. If your upstream fee is dominated by LLM tokens, the same instrument feeds directly into ",[23,2141,2143],{"href":2142},"\u002Fautomating-side-hustle-operations-with-apis\u002Fautomating-ai-workflows-with-python-apis\u002Fcontrolling-llm-api-costs-in-production\u002F","controlling LLM API costs in production",[165,2145],{},[168,2147,2149],{"id":2148},"step-4-turn-metrics-into-alerts-and-slos","Step 4: Turn metrics into alerts and SLOs",[14,2151,2152],{},"A dashboard nobody is watching at 3 a.m. is not monitoring — it is decoration. The point of the histogram and counters is to page you before a customer does. Define one or two service level objectives, express them as PromQL, and let your alerting backend fire when they break. For a small API, two alerts cover most of the ground: a latency SLO on the tail, and an error-budget-style alert on the 5xx rate.",[311,2154,2157],{"className":2155,"code":2156,"language":95,"meta":316},[1815],"# ALERT: p99 latency over 1s for 10 minutes on any route\nhistogram_quantile(0.99, sum(rate(http_request_duration_seconds_bucket[5m])) by (le, route)) > 1\n\n# ALERT: 5xx error rate above 2% over 5 minutes\nsum(rate(http_requests_total{status=~\"5..\"}[5m]))\n  \u002F sum(rate(http_requests_total[5m])) > 0.02\n",[18,2158,2156],{"__ignoreMap":316},[14,2160,2161,2162,2164,2165,2169],{},"Keep the alert set small on purpose. An inbox that fires ten warnings a day trains you to ignore all of them, and the one that mattered arrives buried. Two or three high-signal alerts that only fire when a real customer is affected beat a wall of noisy thresholds. Tie each alert to a symptom a customer would feel — slow responses, failed requests, a stalled webhook backlog — not to an internal number like CPU that may be high for perfectly healthy reasons. When an alert fires, the ",[18,2163,20],{}," in your logs and the route label on the metric together point you straight at the failing calls, and if the slow span is a database round trip, that is your cue to look at ",[23,2166,2168],{"href":2167},"\u002Fscaling-and-operating-production-python-apis\u002Fasync-database-access-with-sqlalchemy\u002F","async database access with SQLAlchemy"," and connection pool sizing.",[165,2171],{},[168,2173,2175],{"id":2174},"step-5-optional-opentelemetry-tracing","Step 5 (optional): OpenTelemetry tracing",[14,2177,2178,2179,2183],{},"Metrics tell you a route is slow; a trace tells you ",[2180,2181,2182],"em",{},"which span"," — the database call, the upstream API, or your own code. When a single request fans out to several services, add OpenTelemetry. It is opt-in and gated on the collector endpoint being set:",[311,2185,2187],{"className":510,"code":2186,"language":512,"meta":316,"style":316},"# tracing.py\nimport os\nfrom opentelemetry import trace\nfrom opentelemetry.sdk.trace import TracerProvider\nfrom opentelemetry.sdk.trace.export import BatchSpanProcessor\nfrom opentelemetry.sdk.resources import Resource\nfrom opentelemetry.exporter.otlp.proto.grpc.trace_exporter import OTLPSpanExporter\nfrom opentelemetry.instrumentation.fastapi import FastAPIInstrumentor\n\ndef configure_tracing(app) -> None:\n    endpoint = os.getenv(\"OTEL_EXPORTER_OTLP_ENDPOINT\")\n    if not endpoint:\n        return  # tracing stays off until an endpoint is configured\n    resource = Resource.create(\n        {\"service.name\": os.getenv(\"SERVICE_NAME\", \"builder-api\")}\n    )\n    provider = TracerProvider(resource=resource)\n    provider.add_span_processor(\n        BatchSpanProcessor(OTLPSpanExporter(endpoint=endpoint))\n    )\n    trace.set_tracer_provider(provider)\n    FastAPIInstrumentor.instrument_app(app)\n",[18,2188,2189,2194,2200,2212,2224,2236,2248,2260,2272,2276,2290,2304,2314,2321,2331,2351,2355,2373,2378,2391,2395,2400],{"__ignoreMap":316},[320,2190,2191],{"class":103,"line":322},[320,2192,2193],{"class":342},"# tracing.py\n",[320,2195,2196,2198],{"class":103,"line":339},[320,2197,525],{"class":524},[320,2199,536],{"class":528},[320,2201,2202,2204,2207,2209],{"class":103,"line":346},[320,2203,787],{"class":524},[320,2205,2206],{"class":528}," opentelemetry ",[320,2208,525],{"class":524},[320,2210,2211],{"class":528}," trace\n",[320,2213,2214,2216,2219,2221],{"class":103,"line":539},[320,2215,787],{"class":524},[320,2217,2218],{"class":528}," opentelemetry.sdk.trace ",[320,2220,525],{"class":524},[320,2222,2223],{"class":528}," TracerProvider\n",[320,2225,2226,2228,2231,2233],{"class":103,"line":547},[320,2227,787],{"class":524},[320,2229,2230],{"class":528}," opentelemetry.sdk.trace.export ",[320,2232,525],{"class":524},[320,2234,2235],{"class":528}," BatchSpanProcessor\n",[320,2237,2238,2240,2243,2245],{"class":103,"line":554},[320,2239,787],{"class":524},[320,2241,2242],{"class":528}," opentelemetry.sdk.resources ",[320,2244,525],{"class":524},[320,2246,2247],{"class":528}," Resource\n",[320,2249,2250,2252,2255,2257],{"class":103,"line":573},[320,2251,787],{"class":524},[320,2253,2254],{"class":528}," opentelemetry.exporter.otlp.proto.grpc.trace_exporter ",[320,2256,525],{"class":524},[320,2258,2259],{"class":528}," OTLPSpanExporter\n",[320,2261,2262,2264,2267,2269],{"class":103,"line":596},[320,2263,787],{"class":524},[320,2265,2266],{"class":528}," opentelemetry.instrumentation.fastapi ",[320,2268,525],{"class":524},[320,2270,2271],{"class":528}," FastAPIInstrumentor\n",[320,2273,2274],{"class":103,"line":626},[320,2275,551],{"emptyLinePlaceholder":550},[320,2277,2278,2280,2283,2286,2288],{"class":103,"line":632},[320,2279,557],{"class":524},[320,2281,2282],{"class":325}," configure_tracing",[320,2284,2285],{"class":528},"(app) -> ",[320,2287,567],{"class":566},[320,2289,570],{"class":528},[320,2291,2292,2295,2297,2299,2302],{"class":103,"line":643},[320,2293,2294],{"class":528},"    endpoint ",[320,2296,579],{"class":524},[320,2298,582],{"class":528},[320,2300,2301],{"class":329},"\"OTEL_EXPORTER_OTLP_ENDPOINT\"",[320,2303,827],{"class":528},[320,2305,2306,2308,2311],{"class":103,"line":649},[320,2307,1536],{"class":524},[320,2309,2310],{"class":524}," not",[320,2312,2313],{"class":528}," endpoint:\n",[320,2315,2316,2318],{"class":103,"line":655},[320,2317,1563],{"class":524},[320,2319,2320],{"class":342},"  # tracing stays off until an endpoint is configured\n",[320,2322,2323,2326,2328],{"class":103,"line":672},[320,2324,2325],{"class":528},"    resource ",[320,2327,579],{"class":524},[320,2329,2330],{"class":528}," Resource.create(\n",[320,2332,2333,2336,2339,2342,2344,2346,2348],{"class":103,"line":678},[320,2334,2335],{"class":528},"        {",[320,2337,2338],{"class":329},"\"service.name\"",[320,2340,2341],{"class":528},": os.getenv(",[320,2343,819],{"class":329},[320,2345,490],{"class":528},[320,2347,824],{"class":329},[320,2349,2350],{"class":528},")}\n",[320,2352,2353],{"class":103,"line":684},[320,2354,724],{"class":528},[320,2356,2357,2360,2362,2365,2368,2370],{"class":103,"line":695},[320,2358,2359],{"class":528},"    provider ",[320,2361,579],{"class":524},[320,2363,2364],{"class":528}," TracerProvider(",[320,2366,2367],{"class":602},"resource",[320,2369,579],{"class":524},[320,2371,2372],{"class":528},"resource)\n",[320,2374,2375],{"class":103,"line":701},[320,2376,2377],{"class":528},"    provider.add_span_processor(\n",[320,2379,2380,2383,2386,2388],{"class":103,"line":707},[320,2381,2382],{"class":528},"        BatchSpanProcessor(OTLPSpanExporter(",[320,2384,2385],{"class":602},"endpoint",[320,2387,579],{"class":524},[320,2389,2390],{"class":528},"endpoint))\n",[320,2392,2393],{"class":103,"line":721},[320,2394,724],{"class":528},[320,2396,2397],{"class":103,"line":727},[320,2398,2399],{"class":528},"    trace.set_tracer_provider(provider)\n",[320,2401,2402],{"class":103,"line":732},[320,2403,2404],{"class":528},"    FastAPIInstrumentor.instrument_app(app)\n",[14,2406,1787,2407,2410],{},[18,2408,2409],{},"BatchSpanProcessor"," exports spans asynchronously off the request path, so it adds negligible latency. Sample aggressively in production — exporting every span at high traffic is both expensive and rarely necessary. A 5–10% head sample keeps a representative view of latency while cutting export volume and your tracing backend's bill by an order of magnitude. The one time you want higher sampling is right after a deploy or during an incident, when you can temporarily raise the rate to catch the specific failure you are chasing, then drop it back.",[165,2412],{},[168,2414,2416],{"id":2415},"configuration-reference","Configuration reference",[379,2418,2419,2431],{},[382,2420,2421],{},[385,2422,2423,2425,2428],{},[388,2424,390],{},[388,2426,2427],{},"Default",[388,2429,2430],{},"Production recommendation",[398,2432,2433,2452,2465,2483,2495],{},[385,2434,2435,2439,2443],{},[403,2436,2437],{},[18,2438,407],{},[403,2440,2441],{},[18,2442,415],{},[403,2444,2445,2447,2448,2451],{},[18,2446,415],{},"; drop to ",[18,2449,2450],{},"WARNING"," only if log volume\u002Fcost forces it",[385,2453,2454,2458,2462],{},[403,2455,2456],{},[18,2457,422],{},[403,2459,2460],{},[18,2461,430],{},[403,2463,2464],{},"Unique per deployable service for clean filtering",[385,2466,2467,2471,2475],{},[403,2468,2469],{},[18,2470,452],{},[403,2472,2473],{},[18,2474,463],{},[403,2476,2477,2479,2480,2482],{},[18,2478,463],{},"; protect ",[18,2481,33],{}," at the network\u002Fproxy layer",[385,2484,2485,2489,2492],{},[403,2486,2487],{},[18,2488,437],{},[403,2490,2491],{},"unset (tracing off)",[403,2493,2494],{},"Set to your collector; enable sampling",[385,2496,2497,2502,2506],{},[403,2498,2499],{},[18,2500,2501],{},"OTEL_TRACES_SAMPLER_ARG",[403,2503,2504],{},[18,2505,1325],{},[403,2507,2508,2510,2511,2513],{},[18,2509,1305],{},"–",[18,2512,1310],{}," at meaningful traffic",[165,2515],{},[168,2517,2519],{"id":2518},"gotchas-and-failure-modes","Gotchas and failure modes",[2521,2522,2523,2534,2540,2546,2559,2569],"ol",{},[2524,2525,2526,2529,2530,2533],"li",{},[176,2527,2528],{},"Logging PII."," A ",[18,2531,2532],{},"log.info(\"login\", email=user.email)"," line ships personal data to a log store you may not control, creating a compliance liability under GDPR or CCPA that can outlast the bug you were debugging. Log a hashed user id, never raw emails, tokens, or card data. Add a structlog processor that redacts known-sensitive keys before the renderer runs, so the redaction is structural rather than something every caller has to remember.",[2524,2535,2536,2539],{},[176,2537,2538],{},"High-cardinality metric labels."," Putting user ids, request paths with ids, or raw query strings in label values creates unbounded time series and can crash Prometheus. Labels must be low-cardinality: method, route template, status, tier — nothing per-user. If you need per-user cost accounting, that belongs in a log line or a database row, not a metric label.",[2524,2541,2542,2545],{},[176,2543,2544],{},"Blocking log I\u002FO."," Writing logs synchronously to a slow sink (a remote HTTP endpoint, an unbuffered file) stalls the event loop and tanks tail latency for every concurrent request sharing that worker. Log to stdout and let the platform ship them; never make a network call inside the request path to emit a log.",[2524,2547,2548,2551,2552,2555,2556,2558],{},[176,2549,2550],{},"Unsampled traces and debug logs in prod."," ",[18,2553,2554],{},"DEBUG"," logging at scale buries signal and inflates your log bill; 100% trace sampling does the same to your tracing backend. Default to ",[18,2557,415],{}," and single-digit-percent sampling.",[2524,2560,2561,2568],{},[176,2562,2563,2564,2567],{},"Counting on ",[18,2565,2566],{},"request.url.path"," for metrics."," As covered in Step 2, this is the fast path to a cardinality explosion. Always label with the route template.",[2524,2570,2571,2574],{},[176,2572,2573],{},"Alerting on averages."," A mean latency alert stays green while your p99 quietly triples, because the tail is a rounding error in the average. Alert on the percentile you actually promise customers, not the mean.",[165,2576],{},[168,2578,2580],{"id":2579},"verification","Verification",[14,2582,2583],{},"Start the service and confirm both signals. The metrics endpoint should return Prometheus text:",[311,2585,2587],{"className":313,"code":2586,"language":315,"meta":316,"style":316},"curl -s http:\u002F\u002Flocalhost:8000\u002Fmetrics | grep http_request_duration_seconds\n# http_request_duration_seconds_bucket{le=\"0.1\",method=\"GET\",route=\"\u002Fwork\",status=\"200\"} 1.0\n# http_request_duration_seconds_count{method=\"GET\",route=\"\u002Fwork\",status=\"200\"} 1.0\n",[18,2588,2589,2609,2614],{"__ignoreMap":316},[320,2590,2591,2594,2597,2600,2603,2606],{"class":103,"line":322},[320,2592,2593],{"class":325},"curl",[320,2595,2596],{"class":566}," -s",[320,2598,2599],{"class":329}," http:\u002F\u002Flocalhost:8000\u002Fmetrics",[320,2601,2602],{"class":524}," |",[320,2604,2605],{"class":325}," grep",[320,2607,2608],{"class":329}," http_request_duration_seconds\n",[320,2610,2611],{"class":103,"line":339},[320,2612,2613],{"class":342},"# http_request_duration_seconds_bucket{le=\"0.1\",method=\"GET\",route=\"\u002Fwork\",status=\"200\"} 1.0\n",[320,2615,2616],{"class":103,"line":346},[320,2617,2618],{"class":342},"# http_request_duration_seconds_count{method=\"GET\",route=\"\u002Fwork\",status=\"200\"} 1.0\n",[14,2620,2621,2622,2624],{},"And a request should emit a single-line JSON log carrying the ",[18,2623,20],{},":",[311,2626,2630],{"className":2627,"code":2628,"language":2629,"meta":316,"style":316},"language-json shiki shiki-themes github-light github-dark","{\"request_id\": \"8f1c...\", \"service\": \"builder-api\", \"path\": \"\u002Fwork\", \"method\": \"GET\", \"status\": 200, \"duration_ms\": 3.41, \"level\": \"info\", \"timestamp\": \"2026-06-18T10:00:00Z\", \"event\": \"request_completed\"}\n","json",[18,2631,2632],{"__ignoreMap":316},[320,2633,2634,2637,2640,2642,2645,2647,2650,2652,2654,2656,2659,2661,2663,2665,2667,2669,2672,2674,2676,2678,2681,2683,2686,2688,2691,2693,2696,2698,2701,2703,2706,2708,2711,2713,2716,2718,2720],{"class":103,"line":322},[320,2635,2636],{"class":528},"{",[320,2638,2639],{"class":566},"\"request_id\"",[320,2641,1779],{"class":528},[320,2643,2644],{"class":329},"\"8f1c...\"",[320,2646,490],{"class":528},[320,2648,2649],{"class":566},"\"service\"",[320,2651,1779],{"class":528},[320,2653,824],{"class":329},[320,2655,490],{"class":528},[320,2657,2658],{"class":566},"\"path\"",[320,2660,1779],{"class":528},[320,2662,1752],{"class":329},[320,2664,490],{"class":528},[320,2666,1271],{"class":566},[320,2668,1779],{"class":528},[320,2670,2671],{"class":329},"\"GET\"",[320,2673,490],{"class":528},[320,2675,1281],{"class":566},[320,2677,1779],{"class":528},[320,2679,2680],{"class":566},"200",[320,2682,490],{"class":528},[320,2684,2685],{"class":566},"\"duration_ms\"",[320,2687,1779],{"class":528},[320,2689,2690],{"class":566},"3.41",[320,2692,490],{"class":528},[320,2694,2695],{"class":566},"\"level\"",[320,2697,1779],{"class":528},[320,2699,2700],{"class":329},"\"info\"",[320,2702,490],{"class":528},[320,2704,2705],{"class":566},"\"timestamp\"",[320,2707,1779],{"class":528},[320,2709,2710],{"class":329},"\"2026-06-18T10:00:00Z\"",[320,2712,490],{"class":528},[320,2714,2715],{"class":566},"\"event\"",[320,2717,1779],{"class":528},[320,2719,1011],{"class":329},[320,2721,1784],{"class":528},[14,2723,2724,2725,2727,2728,2730,2731,2733,2734,2736,2737,471],{},"If the log is multi-line plain text, your formatter is not wired up; if ",[18,2726,33],{}," is empty, the middleware is registered after the route that handles ",[18,2729,33],{},". Fold both checks into your test suite so a regression in the middleware order fails CI rather than production — a request through the app should assert a ",[18,2732,2680],{}," on ",[18,2735,33],{}," and a parseable JSON log line, which fits naturally into ",[23,2738,2740],{"href":2739},"\u002Fscaling-and-operating-production-python-apis\u002Ftesting-python-apis-with-pytest\u002F","testing Python APIs with pytest",[165,2742],{},[168,2744,2746],{"id":2745},"cost-and-performance-note","Cost and performance note",[14,2748,2749],{},"Instrumentation is not free, but the overhead is small and bounded against the alternative of operating blind. A Prometheus histogram observation and an in-process counter increment are microsecond-scale; the JSON log line costs a serialization plus a stdout write. In practice the per-request overhead of logging plus metrics is well under a millisecond — invisible next to a single database round-trip that costs several.",[14,2751,2752,2753,2755,2756,2758,2759,2761,2762,2764,2765,2767],{},"The real cost lever is volume, not per-call work. Put concrete numbers on it: at a steady 50 requests per second you serve roughly 130 million requests a month, and at one ",[18,2754,415],{}," log line of a few hundred bytes each that is on the order of 30–60 GB of logs. Most hosted log backends charge a few dollars per GB ingested, so ",[18,2757,415],{}," logging costs you a manageable double-digit sum, while flipping to ",[18,2760,2554],{}," — which can emit five to ten lines per request — turns that into hundreds of GB and a bill that dwarfs your compute. Metrics scale with series count rather than request volume, so a handful of low-cardinality labels keeps Prometheus in the noise; the moment a stray high-cardinality label creates millions of series, storage and memory costs jump non-linearly. Keep ",[18,2763,407],{}," at ",[18,2766,415],{},", sample traces, and hold your labels low-cardinality, and you get full operational visibility for a rounding error on latency and a predictable, small spend on storage. The blind spot — not knowing your p99 or your cost-per-request until a customer or an invoice tells you — is far more expensive than any of it.",[165,2769],{},[168,2771,2773],{"id":2772},"faq","FAQ",[14,2775,2776,2779,2780,2782,2783,2785],{},[176,2777,2778],{},"How much does monitoring add to my hosting bill at 1M requests a month?","\nAt a million requests a month with one ",[18,2781,415],{}," log line each, expect on the order of half a gigabyte of logs — a few dollars on most hosted backends — plus effectively nothing for metrics, since Prometheus cost scales with series count, not request volume. The overhead only balloons if you leave ",[18,2784,2554],{}," logging on or let a high-cardinality label multiply your metric series; both are configuration mistakes, not inherent costs, and both are avoidable with the defaults in this guide.",[14,2787,2788,2791,2792,2794],{},[176,2789,2790],{},"Do I need both structlog and Prometheus, or is one enough?","\nThey answer different questions and you want both. Structured logs give you per-request detail you can grep by ",[18,2793,20],{}," for debugging; Prometheus metrics give you aggregate trends, percentiles, and alerting across all traffic. Logs without metrics mean no dashboards or alerts; metrics without logs mean you can see a spike but cannot trace a single failing request. Together they cost a few dollars a month at small scale, which is far cheaper than one prolonged outage you debug by guesswork.",[14,2796,2797,2800,2801,2804],{},[176,2798,2799],{},"How do I compute p95 and p99 without writing percentile code?","\nExpose a Prometheus histogram (Step 2) and use ",[18,2802,2803],{},"histogram_quantile(0.95, ...)"," in PromQL. The histogram buckets accumulate in-process; your monitoring backend does the percentile math at query time, so you never compute or store percentiles yourself. The one thing you control is the bucket boundaries — set them to bracket your real target latency so the quantile lands on a tight bucket rather than a wide one, or the number will be imprecise.",[14,2806,2807,2810,2811,2813],{},[176,2808,2809],{},"Should the \u002Fmetrics endpoint be public?","\nNo. It leaks operational detail and route names that help an attacker map your service. Keep it reachable only from your scraper — bind it to an internal network, restrict it at the reverse proxy by IP, or require an auth header. The ",[18,2812,452],{}," flag lets you disable it entirely in environments that should not expose it at all.",[14,2815,2816,2819,2820,2823],{},[176,2817,2818],{},"Does cost-per-request tracking need a separate billing system?","\nNot to observe it. The Prometheus counter in Step 3 gives you average cost-per-request per tier for monitoring and margin alerts. You only need a billing system such as Stripe metered billing when you want to ",[2180,2821,2822],{},"charge"," customers for that usage, which is a separate concern from watching your own unit economics — and if your margins are eroding, you often want to know weeks before you change what you bill.",[165,2825],{},[168,2827,2829],{"id":2828},"related","Related",[14,2831,2832],{},[176,2833,2834],{},"Same section:",[2836,2837,2838,2844,2850,2856],"ul",{},[2524,2839,2840,2843],{},[23,2841,2842],{"href":365},"Structured Logging with structlog"," — the full processor pipeline behind Step 1.",[2524,2845,2846,2849],{},[23,2847,2848],{"href":2739},"Testing Python APIs with pytest"," — assert your middleware order and log schema in CI.",[2524,2851,2852,2855],{},[23,2853,2854],{"href":2167},"Async Database Access with SQLAlchemy"," — where a slow trace span usually leads.",[2524,2857,2858,2860],{},[23,2859,26],{"href":25}," — the parent section overview.",[14,2862,2863],{},[176,2864,2865],{},"Other sections:",[2836,2867,2868,2874],{},[2524,2869,2870,2873],{},[23,2871,2872],{"href":37},"Designing API Pricing Tiers"," — turn cost-per-request into defensible margins.",[2524,2875,2876,2879],{},[23,2877,2878],{"href":2132},"Logging API Usage Events to Postgres"," — persist usage for customer-facing reporting.",[2881,2882,2883],"style",{},"html pre.shiki code .sScJk, html code.shiki .sScJk{--shiki-default:#6F42C1;--shiki-dark:#B392F0}html pre.shiki code .sZZnC, html code.shiki .sZZnC{--shiki-default:#032F62;--shiki-dark:#9ECBFF}html pre.shiki code .sJ8bj, html code.shiki .sJ8bj{--shiki-default:#6A737D;--shiki-dark:#6A737D}html .default .shiki span {color: var(--shiki-default);background: var(--shiki-default-bg);font-style: var(--shiki-default-font-style);font-weight: var(--shiki-default-font-weight);text-decoration: var(--shiki-default-text-decoration);}html .shiki span {color: var(--shiki-default);background: var(--shiki-default-bg);font-style: var(--shiki-default-font-style);font-weight: var(--shiki-default-font-weight);text-decoration: var(--shiki-default-text-decoration);}html .dark .shiki span {color: var(--shiki-dark);background: var(--shiki-dark-bg);font-style: var(--shiki-dark-font-style);font-weight: var(--shiki-dark-font-weight);text-decoration: var(--shiki-dark-text-decoration);}html.dark .shiki span {color: var(--shiki-dark);background: var(--shiki-dark-bg);font-style: var(--shiki-dark-font-style);font-weight: var(--shiki-dark-font-weight);text-decoration: var(--shiki-dark-text-decoration);}html pre.shiki code .szBVR, html code.shiki .szBVR{--shiki-default:#D73A49;--shiki-dark:#F97583}html pre.shiki code .sVt8B, html code.shiki .sVt8B{--shiki-default:#24292E;--shiki-dark:#E1E4E8}html pre.shiki code .sj4cs, html code.shiki .sj4cs{--shiki-default:#005CC5;--shiki-dark:#79B8FF}html pre.shiki code .s4XuR, html code.shiki .s4XuR{--shiki-default:#E36209;--shiki-dark:#FFAB70}",{"title":316,"searchDepth":339,"depth":339,"links":2885},[2886,2887,2888,2889,2890,2891,2892,2893,2894,2895,2896,2897,2898],{"id":170,"depth":339,"text":171},{"id":300,"depth":339,"text":301},{"id":476,"depth":339,"text":477},{"id":1178,"depth":339,"text":1179},{"id":1807,"depth":339,"text":1808},{"id":2148,"depth":339,"text":2149},{"id":2174,"depth":339,"text":2175},{"id":2415,"depth":339,"text":2416},{"id":2518,"depth":339,"text":2519},{"id":2579,"depth":339,"text":2580},{"id":2745,"depth":339,"text":2746},{"id":2772,"depth":339,"text":2773},{"id":2828,"depth":339,"text":2829},"Add JSON structured logging, a Prometheus \u002Fmetrics endpoint, and p95\u002Fp99, error-rate, and cost-per-request tracking to your production Python API.","md",{"pageTitle":5,"datePublished":2902,"dateModified":2903},"2026-06-18","2026-07-23","\u002Fscaling-and-operating-production-python-apis\u002Fmonitoring-and-logging-python-apis",{"title":5,"description":2899},"scaling-and-operating-production-python-apis\u002Fmonitoring-and-logging-python-apis\u002Findex","rVVvo3g-U4w7R9TOIZjbbSrZLNRTvcKaamcTdVvrnNM",{"@context":2909,"@type":2910,"mainEntity":2911},"https:\u002F\u002Fschema.org","FAQPage",[2912,2917,2920,2923,2926],{"@type":2913,"name":2778,"acceptedAnswer":2914},"Question",{"@type":2915,"text":2916},"Answer","At a million requests a month with one INFO log line each, expect on the order of half a gigabyte of logs — a few dollars on most hosted backends — plus effectively nothing for metrics, since Prometheus cost scales with series count, not request volume. The overhead only balloons if you leave DEBUG logging on or let a high-cardinality label multiply your metric series; both are configuration mistakes, not inherent costs, and both are avoidable with the defaults in this guide.",{"@type":2913,"name":2790,"acceptedAnswer":2918},{"@type":2915,"text":2919},"They answer different questions and you want both. Structured logs give you per-request detail you can grep by request_id for debugging; Prometheus metrics give you aggregate trends, percentiles, and alerting across all traffic. Logs without metrics mean no dashboards or alerts; metrics without logs mean you can see a spike but cannot trace a single failing request. Together they cost a few dollars a month at small scale, which is far cheaper than one prolonged outage you debug by guesswork.",{"@type":2913,"name":2799,"acceptedAnswer":2921},{"@type":2915,"text":2922},"Expose a Prometheus histogram (Step 2) and use histogram_quantile(0.95, ...) in PromQL. The histogram buckets accumulate in-process; your monitoring backend does the percentile math at query time, so you never compute or store percentiles yourself. The one thing you control is the bucket boundaries — set them to bracket your real target latency so the quantile lands on a tight bucket rather than a wide one, or the number will be imprecise.",{"@type":2913,"name":2809,"acceptedAnswer":2924},{"@type":2915,"text":2925},"No. It leaks operational detail and route names that help an attacker map your service. Keep it reachable only from your scraper — bind it to an internal network, restrict it at the reverse proxy by IP, or require an auth header. The METRICS_ENABLED flag lets you disable it entirely in environments that should not expose it at all.",{"@type":2913,"name":2818,"acceptedAnswer":2927},{"@type":2915,"text":2928},"Not to observe it. The Prometheus counter in Step 3 gives you average cost-per-request per tier for monitoring and margin alerts. You only need a billing system such as Stripe metered billing when you want to charge customers for that usage, which is a separate concern from watching your own unit economics — and if your margins are eroding, you often want to know weeks before you change what you bill. ---",1784887027769]