Skip to content
THE NYCG / AI

AI development, model training & evaluation

Turn a repeated task, a collection of documents or a model experiment into a defined engineering project. Start with the outcome; we can work through the technology together.

Discuss your AI project

What we can help with.

Explore the work and the deliverables we can scope together.

Agents & workflow automation

Connect an AI assistant to knowledge and business tools so it can help complete a task, with defined permissions and a path to a person.

Possible deliverables
An agreed workflow, connected tools and test cases for completed actions, failures and handoffs.
Technical scope
Work to scope
Knowledge retrieval, tool and MCP integrations, CRM connections, action approvals, failure handling and conversation testing.

Model training & adaptation

Develop or adapt a model for your data and the task it needs to perform.

Possible deliverables
Versioned training work and a comparison on a representative test set.
Technical scope
Work to scope
Large and small language models, embeddings, OCR and computer vision. Data preparation, fine-tuning and comparisons against a baseline.

Evaluation & improvement

Find where an existing model, search system or AI conversation gets things wrong, then test changes.

Possible deliverables
A repeatable evaluation suite, findings and a prioritized improvement plan.
Technical scope
Work to scope
Model, RAG, chat and agent evaluations; source relevance, answer quality, tool behavior and regression testing.

Deployment & model optimization

Prepare a model for the environment where it needs to run, from cloud infrastructure to an embedded device.

Possible deliverables
A deployment approach and measured comparisons against the original model on the target hardware.
Technical scope
Work to scope
Inference deployment, quantization, memory and latency measurement, and hardware-specific validation.
Explore your industry context

Your industry shapes the data, permissions and review process. These examples explain possible approaches; they are not client case studies.

Built by NYCG

From a conversation to useful work.

Our team has built platforms for recruiting automation and AI calling. Explore how they connect context, actions and human review.

Recruiting automation AI calling
Product names, logos and websites
Model training & evaluation

Develop the model.
Test it against the task.

Explore the data, engineering and evaluation involved. We agree the deliverables and acceptance criteria around your workload.

01

Train & adapt

Language, documents and vision.

02

Optimize & deploy

The right model for your hardware.

03

Evaluate

The whole experience, under test.

Build the learning process around your data and the task it needs to perform.

The starting point
A defined task, usable training data and a held-out test set.
The engineering
Dataset preparation, labeling strategy, training pipelines and reproducible experiments.
What we measure
Task accuracy, class-level errors, data leakage and performance on unseen examples.
Deliverables to scope

Versioned data preparation, training runs and a comparison on the test set.

Discuss Model training

Choose and adapt a model to the work, the data boundary and the operating budget.

The starting point
Representative tasks, example responses and deployment constraints.
The engineering
Base-model comparison, fine-tuning or distillation where appropriate, and private deployment.
What we measure
Task success, instruction following, latency, memory use and cost per completed task.
Deliverables to scope

A model comparison, adaptation artifacts and a deployment plan for the chosen environment.

Discuss Large & small language models

Make search understand the language and distinctions in your own documents.

The starting point
Your corpus, real search queries and examples of relevant results.
The engineering
Embedding selection or fine-tuning, chunking, indexing and reranking.
What we measure
Retrieval recall, ranking quality, multilingual coverage and index size.
Deliverables to scope

A retrieval pipeline and a comparison of search results against labeled queries.

Discuss Embedding models

Turn difficult documents into useful fields, with a path back to the original.

The starting point
Representative scans, layouts, handwriting and fields to extract.
The engineering
OCR model adaptation, layout understanding, extraction and low-confidence review.
What we measure
Character and field accuracy, table extraction and failure rates by document type.
Deliverables to scope

An extraction pipeline, field-level test results and a route for uncertain documents to be reviewed.

Discuss OCR & document models

Teach a model what matters in your cameras, products and operating conditions.

The starting point
Labeled images or video covering real conditions and edge cases.
The engineering
Ultralytics and custom training for detection, segmentation and tracking; export to target hardware.
What we measure
Per-class precision and recall, missed detections and performance on the actual device.
Deliverables to scope

Trained model artifacts, error analysis and an export tested on the target device.

Discuss Computer vision models

Make the model fit the hardware without losing the quality the task requires.

The starting point
A baseline model, calibration data and target CPU, GPU or NPU.
The engineering
Precision reduction, quantization, compilation and inference-runtime tuning.
What we measure
Before-and-after accuracy, response time, throughput, memory and power where measurable.
Deliverables to scope

An optimized model artifact and a repeatable comparison against the original model.

Discuss Optimization & quantization

Know where a model works, where it fails and whether a change is an improvement.

The starting point
A versioned test set and acceptance criteria tied to the task.
The engineering
Model comparisons, error analysis, domain-specific tests and human review.
What we measure
Quality by task and subgroup, reliability, latency and release-to-release regressions.
Deliverables to scope

A repeatable evaluation suite, comparison report and examples of failures to investigate.

Discuss Model evaluation

Test the search and the answer separately, then evaluate the complete experience.

The starting point
Real questions, expected sources and access rules.
The engineering
Retrieval tests, source-grounding checks, citation verification and permission-boundary tests.
What we measure
Retrieval recall, answer faithfulness, citation correctness, abstention and unauthorized disclosure.
Deliverables to scope

Separate retrieval and answer evaluations, failure examples and tests for the next release.

Discuss RAG evaluation

Test whether a conversation actually completes the right work.

The starting point
Multi-turn conversations, tool contracts and escalation scenarios.
The engineering
Conversation simulations, tool-call checks, privacy tests and human-handoff review.
What we measure
Task completion, context retention, correct tool use, recovery and response latency.
Deliverables to scope

A scenario-based test suite and a report of completed tasks, failed actions and handoff behavior.

Discuss Chat & agent evaluation
Before we start

Your questions, answered.

Plan an AI workflow: context, actions and human handoffs
When should we consider custom AI development?

Start with a specific task that existing tools do not handle well. Bring an example of the input, the result you need and what goes wrong today. We can help assess whether configuration, an integration or custom development is the useful next step.

Can you improve an existing RAG or AI chat system?

Yes. We can scope an evaluation around representative questions, retrieved sources and expected answers. For an agent, that also includes tool actions, failed requests and human handoffs. The starting point is a repeatable baseline so changes can be compared.

Can an AI agent connect to our CRM through MCP or APIs?

MCP and API integrations are part of our engineering scope. We first establish which records an agent may read, which actions it may request and who approves sensitive changes. Compatibility, access and failure handling are checked for the systems involved.

What do you need before discussing model training?

Describe the task, the model if you already have one, the type of data and where it needs to run. Include known quality, latency or memory constraints. A description is enough for an initial enquiry; access to private data can be agreed separately.

Start a conversation

Bring the problem.
We’ll work through it.

A rough description is enough. If you have technical requirements, bring those too.

Discuss your AI projectBrowse platforms & integrations