What do TensorFlow development services include?
A TensorFlow project is more than a trained model. It covers the data pipeline that feeds training, the Keras model itself, the TFX pipeline that validates data, retrains and evaluates each new version, and the serving layer that answers requests: TF Serving behind an API, LiteRT inside a mobile or edge app, or TensorFlow.js in the browser.
At Miracuves the handoff is the repository with its commit history, the SavedModel (and LiteRT file, where the model runs on a device), the TFX pipeline definition, TensorBoard logs, Grafana dashboards for latency and drift, and a runbook for retraining and rollback. A scoped module covering one use case starts from $3,699 and ships in 4-8 weeks.
TensorFlow, PyTorch or JAX - which should you build on?
For most products the framework does not decide accuracy; your data does. It decides how the model is served, where it can run and who can maintain it. Our honest reading:
- TensorFlow: the strongest end-to-end deployment story from one vendor - TFX pipelines, TF Serving, LiteRT on Android, iOS and microcontrollers, and TensorFlow.js in the browser.
- PyTorch: the default for much of today's research and for many open models, so it is often the faster route when you are adapting a new published model.
- JAX: Google's high-performance research framework, strong on TPUs and very large training runs, with a smaller production toolchain around it.
- Keras 3 runs on all three backends, so a Keras model is not locked to TensorFlow if you change your mind later.
What changed with Keras 3 and LiteRT, and does it affect my project?
Two naming changes matter to anyone reading older TensorFlow tutorials. Keras 3 became multi-backend: the same Keras code can run on TensorFlow, JAX or PyTorch, and current TensorFlow releases ship it as their Keras. Some Keras 2 habits now raise warnings or need small changes, so older model code usually needs a short upgrade pass.
TensorFlow Lite was renamed LiteRT by Google in 2024. Existing .tflite models keep working, and Google also provides tools to convert PyTorch and JAX models into it. For you this means on-device deployment is no longer a reason to pick TensorFlow alone - but a Keras model still has the shortest path to a phone.
When should a TensorFlow model run on the device instead of a server?
Run it on the device when the answer has to come back instantly, work offline, or keep images, audio or health data on the phone - camera-based defect checks, on-device text classification, keyword spotting. We convert the Keras model to LiteRT and quantize it, usually to 8-bit integers, so it is smaller and faster, then check that accuracy after quantization still meets the target agreed at scoping.
Keep it on a server with TF Serving when the model is large, changes often, or needs data the device does not have, such as a recommendation model that reads a live catalogue. Many products use both: a small LiteRT model for the first pass and a larger served model for the hard cases.
How are TensorFlow models served in production, and what does it cost to run?
TF Serving loads a SavedModel, exposes it over REST and gRPC, batches requests, and swaps in a new model version without downtime; we run it on Kubernetes or on a managed platform such as Google Cloud Vertex AI or AWS SageMaker. Because each version sits in its own numbered folder, rolling back is a configuration change.
Running cost is driven by one choice more than any other: whether inference needs a GPU. Many tabular, forecasting and ranking models serve comfortably on CPUs; large vision and sequence models may not. We profile latency and throughput on your expected traffic before launch, so the monthly hosting figure is known before you commit.
Can Miracuves take over or migrate an existing TensorFlow codebase?
Yes. Most inherited TensorFlow work falls into three groups: TensorFlow 1.x code built on sessions, placeholders and Estimators, which Google has deprecated; Keras 2 models that need updating for Keras 3; and notebooks with no pipeline around them. We start with a written audit of the code, the data and the serving path, then migrate in steps so the production model keeps answering while the new pipeline is built.
If your team would rather be on PyTorch, or is moving from PyTorch to TensorFlow for deployment, we will say whether a Keras 3 port, a conversion or a retrain is cheapest. Migrations are quoted as custom work before they start.
Do you need a TensorFlow specialist, or a broader AI team?
Hire a TensorFlow team when the model must be served at scale, run on a device or retrain on a schedule - that is where TFX, TF Serving and LiteRT pay off. If you are still choosing between classical models and deep learning, our ML development service starts one step earlier. If the product is built around a large language model, LLM development is the better fit, and image-first products usually belong with computer vision development.
Compared with hiring TensorFlow engineers yourself, a company build gives you the pipeline, serving and monitoring skills in one team from the first week. When the model is live, you can keep us on a retainer from $2,299/month or hand the repository to your own ML team.