FDE PulseFDE jobs open 441New in 7 days 29Companies hiring 47Remote-friendly 24%Median US pay $216kTop hirer Databricks 125
VI

The newspaper of the Forward Deployed Engineer

Tools

Modal: from a Python function to a GPU demo and a batch job in one afternoon

When a client asks "can we try it on real data?", an FDE rarely has time to build infrastructure. Modal lets you skip most of that work.

Modal: from a Python function to a GPU demo and a batch job in one afternoon
Photo: Raysonho @ Open Grid Scheduler / Grid Engine / CC BY 3.0

In brief

  • Modal runs your Python function in a cloud container, bills by the second and charges nothing while resources sit idle
  • modal serve gives you a demo URL that updates as you edit code, .map spreads work across many containers, and one decorator parameter gets you a GPU
  • Modal is strong for batch jobs and prototypes; for always-on, high-traffic services, consider other options
ShareLinkedInFacebookX
Horizontal bar chart: processing 10,000 contract PDFs at 5 seconds each. Running sequentially on a laptop takes 50,000 seconds (nearly 14 hours); using Modal .map across 100 containers takes about 500 seconds (just over eight minutes). The figure of 100 containers is illustrative only.
Illustrative calculation from the article: the same extract function run sequentially on a laptop versus split across 100 parallel containers with .map, excluding container start-up and model loading time. Source: the article's illustrative calculation.

What sinks a client demo is rarely the model. More often the GPU machine has not been provisioned yet, the container configuration is still broken, and the job that has to process 10,000 files is still running sequentially on your laptop.

Modal targets exactly that work. Modal Labs describes it as an AI infrastructure platform: it takes your code, puts it in a container and runs it in the cloud.

For an FDE, the appeal is the pricing model. The company’s pricing page says plainly that you are billed by the second and never pay for idle resources. The Starter plan includes $30 of free credit a month, enough to practise on your own and build prototypes.

Why does “just Python” matter?

On PyPI, the modal client library is described as a way to access serverless compute on demand from a Python script on your own machine. The library is open source under the Apache-2.0 licence. To get started, you install it with pip, run modal setup to authenticate, then modal run a Python file.

An independent comparison by Introl notes that deploying on Modal needs no YAML or Dockerfile. The same piece argues that Modal suits Python-centric teams that need fast iteration and value developer ergonomics over squeezing costs as low as possible.

On a client site, that is close to the FDE job description: you do not switch languages, you do not need the infrastructure team, and logic you have just tested in a notebook can be pulled out into a function and sent to the cloud within the same session.

A thought experiment: 10,000 contracts and one afternoon

Suppose an insurance company hands you 10,000 PDF contracts and asks whether your clause-extraction model can be trusted. If each file takes 5 seconds, running them sequentially on a laptop takes 50,000 seconds, nearly 14 hours. The afternoon is gone and you still have no results.

On Modal, you keep the extract function as it is, add a decorator and call .map. A minimal sketch looks like this:

import modal

app = modal.App("contract-extraction")

@app.function(gpu="L4")  # choose the GPU type via the gpu parameter
def extract(pdf: bytes) -> dict:
    # load the model, read one contract, return the clauses as JSON
    ...

@app.function()
@modal.fastapi_endpoint()  # one line turns the function into a web endpoint
def try_it(text: str) -> dict:
    return {"received": text}  # where the client can try it out

@app.local_entrypoint()
def main():
    files = [...]  # contents of the 10,000 PDF files
    for result in extract.map(files):
        print(result)

You run the batch job with modal run contracts.py, while modal serve contracts.py gives you a temporary URL for the client to try. According to the batch processing documentation, .map pushes the heavy computation to powerful machines and collects the results for you. Modal also provisions the infrastructure itself and can scale to thousands of containers in parallel.

Now redo the arithmetic. Suppose .map splits the 10,000 files across 100 containers running at once: each container gets 100 files, about 500 seconds, a little over eight minutes. The figure of 100 is only illustrative; add the time to start containers and load the model, and you still have results the same afternoon rather than the next morning.

The gpu parameter accepts many types, from T4, L4, A10 and L40S to A100, H100, H200 and B200. Practical advice: start with the smallest type that can run the model, and move up only when you measure that it is too slow.

While you tune the endpoint before a meeting, modal serve updates the app every time the file changes, so you can adjust a prompt or the JSON format without redeploying. Send that URL to the business lead and they can try it straight away on their own contracts.

Per-second billing changes how you run demos

Do the sums: a demo endpoint stays up for five working days, but the client calls it for about 20 minutes in total. With a GPU machine rented on a fixed basis, you pay for all five days. With a model that charges nothing while idle, you pay only for the compute that actually runs.

The price is cold starts. Modal’s documentation defines a cold start as the higher latency incurred when a new container has to start, and for interactive workloads the only remedy is to keep extra containers warm. Sam McKay writes on Enterprise DNA that Modal handles cold starts better than most serverless platforms but does not make them disappear.

Do not let the first slow call land just as the client is watching the screen.

When should Modal stay on the bench?

The Enterprise DNA review is blunt: Modal is strong at batch jobs, but if you force it into an always-on, high-traffic service, cost becomes a burden. The reason lies in the pricing model itself: per-second billing pays off when load comes in bursts and then rests, and loses its advantage when the machine never rests.

So Modal fits best the work an FDE does before the client decides: proving feasibility, running one-off jobs, building prototypes. When a prototype becomes a system that runs all day, recalculate the cost before recommending that the architecture stays as it is.

The remaining limit is not technical. Code runs in cloud containers, so the client’s data has to leave their environment. With banks, hospitals or clients under strict data-residency rules, ask about the policy first, or demo with anonymised data.

What to learn first, and what to put on your CV

Learn in the order of real work: modal run a function, then modal serve to get a URL, then .map over a real dataset, and only then GPUs and measuring cold starts.

The skill most worth practising is turning notebook code into pure functions that take one input and return one output, because .map only pays off when functions are written that way.

When writing a CV or answering FDE interview questions, do not just write “knows Modal”. The advice is to describe an outcome: how many files you processed, how far you cut the run time, which demo URL got the client to say yes.

Fast infrastructure does not make your demo better. It just gives you more time to understand the client’s problem properly.

Was this article useful?

Use with your AI assistantAsk Claude ↗Ask ChatGPT ↗
9 sources
Read next on the roadmap · Stage 5: DeploymentAzure AI Foundry is now Microsoft Foundry: what FDEs need to know when deploying models and agents on a client's AzureAt least one FDE job ad still uses the old name. The hard parts are the ones few people mention: where data is processed, how many RU/s Cosmos DB gets and who creates the private endpoints.