FDE PulseFDE jobs open 434New in the last 7 days 27
VI

The newspaper of the Forward Deployed Engineer

Analysis

LangSmith, Arize, Helicone or PostHog: let the client's constraints choose the tool

A feature matrix tells you what each tool can do. For a forward deployed engineer, two other questions matter more: where the client's data is allowed to go, and what the client actually needs to find out.

In brief

  • Choose the tool by the client's constraints, not by a feature matrix. The first constraint is always where the data is allowed to live.
  • Helicone is the fastest to integrate because you only change the base URL, but your traffic then passes through Helicone's gateway. LangSmith offers a self-hosted version, and Arize has open-source Phoenix for clients that cannot use SaaS.
  • For agents, a trace must record every reasoning step and every tool call. If you log only the final output, you will not know which step broke when the agent fails.
ShareLinkedInFacebookX

To add Helicone to an application built on an OpenAI-compatible SDK, you change exactly one line: point the base URL at ai-gateway.helicone.ai. Integration that fast is tempting. But that same line also means every prompt and every response from the client now passes through a gateway the client does not control.

That is why the question “should we use LangSmith, Arize, Helicone or PostHog?” is usually aimed at the wrong target. All four tools can record LLM calls. Where they differ is in the kind of client each one quietly assumes: where that client lets data go, how much code they are willing to change, and what they are worried about.

This is exactly the FDE’s job. You are not writing a tool review. You are sitting in the client’s meeting room and have to decide within a few sessions. Read the constraints correctly and you pick the right tool first time. Read them wrongly and three weeks later you will be ripping it out and starting again, just as the client was beginning to trust you.

Why can’t you monitor by hand?

IBM defines LLM observability as the collection of metrics, traces and logs from LLM applications. It also argues that manual monitoring is labour-intensive, error-prone and does not scale.

Picture a pilot where the whole team opens log files and reads them line by line. With few requests it works; in production it breaks down exactly as IBM warns.

JetBrains offers a useful distinction. Evaluation answers whether an agent can do the job. Observability answers whether the agent is doing the job, and clients are usually worried about only one of the two.

JetBrains also stresses that for agents, a trace has to follow the line of reasoning through each step, each tool call and each observation; recording only the final output is not enough. Imagine an order-lookup agent with six steps, where the third step calls the wrong tool.

If you log only the final answer, all you know is that the customer received a wrong answer, not why.

Where is the client’s data allowed to live?

This question eliminates more options than any other, so ask it first. A bank or a hospital may not allow prompts containing customer information to leave its infrastructure. An e-commerce start-up may barely care.

LangSmith lets you choose between cloud, hybrid and self-hosted deployment. For clients with data-residency requirements, that is a significant advantage.

Arize solves the problem differently. Arize AX is the commercial platform, and the AX documentation itself links to Phoenix, Arize’s open-source product. If the client cannot use SaaS, Phoenix is worth trying before concluding that you have to build your own tooling.

Helicone is the opposite case. Its quick start places Helicone in the middle of the request path as a gateway, bundled with logging, observability and fallback between providers. With a restricted client, your first task is to ask the security team whether they will accept an external gateway; only then does fast integration enter the conversation.

Gateway or instrumentation: where does the tool sit?

Gateways and instrumentation have almost opposite strengths and weaknesses. A gateway needs almost no code changes, just a new base URL, but it sits directly in the path of every request. Instrumentation requires you to touch more code, but it does not stand between the application and the provider.

Do not assume instrumentation is automatically safer. Traces sent to LangSmith cloud or Arize SaaS still contain prompts and outputs, much as PostHog describes recording the full conversation context for each generation. The data takes a different route, but no less of it leaves the client’s infrastructure.

Arize AX offers auto-instrumentation for more than 30 providers and frameworks, which cuts the code changes considerably. That matters when you meet a codebase mixing several frameworks, which is common at clients that have run a few rounds of experiments before you arrive.

Helicone’s gateway also brings something instrumentation does not: fallback between providers. According to the quick start, Helicone’s credit model passes through provider pricing at “0% markup”, and clients can also bring their own provider keys. For a start-up that wants logs quickly and a backup when one provider fails, it is a pragmatic choice.

What question is the client trying to answer?

This is where JetBrains’ line between evaluation and observability becomes concrete. Arize describes AX as an AI engineering platform for tracing, evaluating and running experiments, which is broader than a logging tool. LangSmith talks about visibility across the whole application, from a single trace to production-wide metrics, with dashboards and alerts.

PostHog takes a different direction. Each LLM call is recorded as a generation, including the full conversation context, tool calls, token counts and latency, and cost is calculated automatically from model pricing. This data sits alongside product analytics, so you can connect user behaviour with model behaviour.

Imagine an e-commerce client already using PostHog to measure funnels. They are not asking whether the agent reasoned correctly at step four. They want to know whether users who chat with the bot buy more, and how many tokens each order costs. For this client, adding a second tool may simply fragment the data.

Client constraint Tool to try first What to check on site
Data must not leave the infrastructure LangSmith self-hosted or Phoenix Does the client’s operations team have the people to run and maintain it?
Wants logs within a single session, minimal code changes Helicone gateway Will the security team accept traffic passing through a gateway?
Codebase mixes several frameworks Arize AX with auto-instrumentation Are the client’s frameworks among the more than 30 supported?
Needs evaluation and experiments, not just logging Arize AX Does the client really need to know whether the agent can do the job, or only whether it is running?
Already uses PostHog, cares about user behaviour and cost PostHog LLM analytics Will traces stay detailed enough step by step as the agent grows more complex?

What should a discovery session ask?

The order of the questions matters as much as their content. Ask about data first, because the answer can eliminate half the options at once. Then ask how much code the client is willing to change, and only then what they want to measure: whether the agent can do the job, whether it is doing the job, or whether users benefit.

A common trap is choosing the most powerful tool for a question the client never asked. A full evaluation and experimentation platform becomes a burden if the client only needs an alert when token costs spike.

Conversely, a simple logging gateway will not be enough when the client is struggling with a multi-step agent that fails somewhere nobody can see.

In regulated markets, the data question comes first

In Vietnam, as in many markets, if your client works in finance, telecoms or the public sector, “where does the data live?” is likely to be asked before anything else. In that setting, someone who has actually run a self-hosted observability tool has a clear advantage over someone who has only used the SaaS version.

On your CV, do not write “experienced with LangSmith”. Frame it as a decision: which tool you chose, because of which constraint, and what failure it helped you find.

When reading an FDE job description, look for phrases such as “on-prem”, “data residency” or “agent reliability”. They tell you which of the three constraints that company’s clients are running into.

These four tools will change their features, change their names, perhaps even merge. But the three questions about data, position in the system and what the client needs to know will not change. Whoever knows to ask those three questions will pick the right tool, including tools that do not exist until next year.

6 sources
Read next on the roadmap · Stage 7: LeadershipHow to turn a finished deployment into a playbook, runbooks and a template repo for the next customerWhat a team learns on a deployment usually leaves with the FDE when the project ends. This guide shows how to keep it in five documents and a template repo.