Senior Data Engineer·Technical Lead

I build the data infrastructure that recognises revenue at scale.

Microsoft Fabric · Azure · Power BI · Databricks · Spark · Python · SQL
Batch and streaming pipelines, governed reporting platforms and AI-driven analytics across banking, telecoms and government.

Ten years owning delivery end to end with client discovery and scoping, architecture and build, then the team that runs it.

Remote · UTC−5 · International
Gilroy Gordon
0Revenue & Cost Impact / month
0Community Members Impacted
0Integrations
0Insight delivery cost cut
0Years building & leading
Business Impact

Data engineering should change
what the business can do.

I build data systems around real business decisions — improving revenue, understanding customers, and making operations more efficient. The engineering matters, but the outcome is what counts.

Grow and Recover Revenue

Turn data into revenue

Sales forecasting, lead generation, marketing ROI and e-commerce analytics that help teams identify where revenue is coming from, where it is leaking, and where to invest next.

Understand Customers

Make customer behaviour useful

Segmentation, behavioural analysis, loyalty and automated customer experiences that turn fragmented customer data into better targeting, retention and engagement.

Run Efficiently

Make the business run smarter

Financial processing, inventory, operational workflows and APIs that reduce manual work, improve visibility and give teams reliable data for faster decisions.

Focus business outcomes Bridge data → decisions Measure impact
High Throughput

Selected work

Client names and system names are withheld under NDA. The architecture, the trade-offs and the results are not.

Telecoms · Streaming · Observability

Streaming Delivery Observability

Event-driven telemetry platform for a high-volume messaging service used by retail banks. Idempotent webhook ingestion with schema validation at the boundary, micro-batched to object storage, status-reconciled in Spark and served with sub-ten-minute freshness. Gave operations the SLA evidence trail that provider credit claims had previously been argued without.

Latency <10 min Volume millions/day Providers multi Mode streaming

StackDockerised Spark · Azure Data Lake · Azure Blob Storage · SQL Server · Python scheduler · Power BI · webhook ingest with schema validation and IP allowlisting

Context
Delivery-status webhooks arriving continuously at high volume. Operations needed per-message status visibility within ten minutes to diagnose failures against provider SLAs.
Options
Managed message bus · managed event streaming service · containerised Spark with batched writes to object storage.
Decision
Containerised Spark, batching webhook payloads to object storage before processing.
Why
Write-cost per event dominated at projected volume; both managed options priced per message. Batching traded roughly two minutes of latency for a large reduction in storage transaction cost, comfortably inside the ten-minute requirement.
Consequence
Materially cheaper and simpler to operate, but the latency floor is now batch-interval-bound. Revisit if the requirement tightens below five minutes.

Business impactConverted an unmeasured operational risk into a contractual position, and made provider performance arguable with evidence.

Reported as generic. Client, platform and provider names withheld.

Finance · Governance · Access control

Governed Financial Reporting Platform

Single governed semantic model over consolidated P&L, with row-level security bound to the organisational hierarchy, provisioning executed as an auditable workflow rather than administrator memory, and lineage documented from source ledger through to executive visual. Replaced a monthly spreadsheet distribution that no auditor could trace.

Access RBAC + RLS Audit full trail Scope all business units Cadence monthly / quarterly

StackSQL Server · automated ETL · Power BI with row-level security · Microsoft Forms provisioning workflow · Python access-management scripts

Context
Directors needed self-serve access to unit-level financials. Finance needed certainty that no leader could see another unit's numbers, and auditors needed evidence of that.
Options
Separate report per business unit · application-layer filtering · row-level security on a single governed model.
Decision
Single semantic model with row-level security, provisioning routed through a tracked request workflow rather than ad-hoc administration.
Why
Per-unit reports would have multiplied maintenance and let definitions drift apart. A single governed model keeps one definition of every metric and makes access an auditable data point rather than an administrator's memory.
Consequence
One model to maintain and one definition of truth, at the cost of a stricter change process — any model change now requires re-validating security roles.

Business impactDirectors self-serve unit financials without a controls exception, and month-end reporting effort collapsed to a refresh.

Reported as generic. Client and system names withheld.

Finance · Reconciliation · Revenue assurance

Revenue Leakage Detection

Entity resolution across two disconnected CRM estates with no shared key. Deterministic matching on normalised identifiers, fuzzy fallback on address and asset serial, and a human-in-the-loop explorer so analysts adjudicate the residual rather than the whole set. Surfaces delivered-but-uninvoiced work orders every close cycle.

Sources 2 CRMs + manual Recovery monthly, recurring Matching automated + assisted

StackSQL Server · ETL pipelines · fuzzy and rule-based matching · Power BI matching explorer with analyst override

Business impactRecurring monthly recovery of revenue already earned and delivered. Figures available on request.

Reported as generic. Client and system names withheld.

Platform · Reliability · Cost

Platform Reliability Monitoring

Observability layer for the reporting estate. Gateway health, dataset refresh outcomes and agent job state ingested into a modelled serving layer, with alerting bound to SLO breach rather than raw failure count. Paired with a capacity scheduler that parks idle compute, so reliability and unit cost improved on the same change.

Failures −95% Cloud spend −30% Detection proactive

StackPower BI Gateway telemetry · SQL Server Agent · Python ETL · custom capacity scheduler pausing idle compute

Context
Refresh failures were discovered by users, not engineers. Separately, development database spend was growing without clear attribution to teams.
Options
Vendor monitoring add-on · manual audit cadence · build telemetry ingestion and a capacity scheduler in-house.
Decision
In-house telemetry pipeline plus a scheduler that pauses capacity when idle.
Why
Availability ranked above cost, but the vendor option priced per-workspace and would have grown with the estate. Building it kept monitoring coverage complete while the scheduler paid for the effort within months.
Consequence
Failures fell sharply and spend dropped roughly a third, but the telemetry pipeline is now itself a system requiring maintenance and its own monitoring.

Business impactFailures found by engineers instead of executives, and a material reduction in platform run-rate on the same piece of work.

Reported as generic. Client and platform names withheld.

Context

The Engineer Behind the Outcomes

The Senior Data Engineer who connects technology to the decisions and outcomes that matter. Explore the work first, then get to know the engineer, experience and capabilities behind it.

Business outcome → enabling technology
Large nodes are business outcomes · small nodes are the stack

Gilroy Gordon is a Data & AI leader, analytics engineer and educator with over ten years delivering measurable impact across private sector, government and nonprofit organisations internationally. He specialises in building scalable data platforms, AI-driven solutions and intelligent automation that improve revenue performance, customer experience and operational efficiency.

A multi-certified Data & Analytics professional, including Microsoft, and a Certified Scrum Master, he combines deep technical expertise with strategic leadership to help organisations translate data into competitive advantage. Gilroy has chaired AI and Data committees, developed Data and AI academic programmes at the University of Technology, Jamaica, and presented at regional and international forums including IEEE workshops and BizTech.

He remains committed to community service, mentorship and capacity building, and has been active in service organisations supporting initiatives that strengthen digital and social impact ecosystems. He is humbled to have been recognised in the Nova Scotia Legislature for contributions to Legal Aid initiatives affecting over 400,000 Canadians, to have been named Young Entrepreneur of the Year in 2020, and to keep finding opportunities in youth development.

The full record

Career history, certifications, education, publications, awards and recommendations are maintained on LinkedIn rather than duplicated here.

Learn more on LinkedIn →

Case Studies

Read about the work, the decisions behind it, and what it delivered.

View Case Studies →
Capabilities

What I work across

Layered the way a platform actually is, rather than as an inventory.

Ingest
Batch and streaming pipelines · CDC · webhook and API ingestion · schema validation
Fabric · ADF · Spark · Kafka · Airflow
Store
Lakehouse and warehouse modelling · partitioning · schema evolution · replayable raw layers
OneLake · Azure Data Lake · SQL Server · Snowflake
Transform
ELT/ETL · modelling · analytics engineering · entity resolution · data quality testing
dbt · PySpark · T-SQL · Databricks · Python
Serve
Semantic models · DAX and performance tuning · executive reporting · embedded analytics · self-serve access
Power BI · Fabric · Tableau · REST APIs
Govern
RBAC and row-level security · lineage · audit trails · access provisioning workflows
Purview · dbt docs · RLS · Fabric domains
Operate
CI/CD for data · observability · incident diagnosis · cost optimisation · mentoring
Azure DevOps · Docker · Git
Everything else in the toolkit — languages, servers, platforms, methods
Languages

SQL · T-SQL · HiveQL · Python · C# · ASP.NET · Java · JavaScript · TypeScript · PHP · R · C · C++ · Dart (Flutter) · Solidity · Cypher · Pig Latin · Bash · HTML5 · CSS

Data & big data

Microsoft Fabric · Azure Data Factory · Azure Databricks · Azure Data Lake · OneLake · Apache Spark · Hadoop · Hive · HBase · Sqoop · Apache Pig · Impala · Zeppelin · Ambari · Kafka · Airflow · dbt · SSIS · Snowflake · KNIME

Databases

SQL Server · MySQL · PostgreSQL · SQLite · MongoDB · Neo4J · Lucene · Solr

BI & reporting

Power BI (Embedded · RLS · Gateway administration · DAX tuning) · SSAS · SSRS · Tableau · Looker Studio · SAP Crystal Reports · Jupyter · R Studio · Excel & Power Query

AI, ML & automation

MLflow · Hugging Face · LLMOps · NLP · Generative AI · recommendation systems · predictive modelling · RPA · Power Automate · web mining

Ship & run

Azure DevOps · Git · GitHub · GitLab · Bitbucket · Team Foundation Server · Subversion · Travis · Docker · Kubernetes · Ansible · Linux · Windows · Android · ARMv6

Application & servers

Node.js · React · Angular · Laravel · Bootstrap · IIS · Apache · NGINX · Tomcat · IBM WebSphere · Glassfish · WordPress · Joomla · Magento

Delivery

Agile · Scrum (Certified ScrumMaster) · Kanban · Waterfall · RAD · Lean Six Sigma (White Belt) · Project Management Essentials

Want to see sample projects including some of these? Filter the archive by technology →
Consumers

Working on something that needs to be trusted?

Two doors, depending on what you need.

Roles & hiring

Talk to me directly

Full professional history, current work and experience live on LinkedIn. Best route for permanent, contract and lead engineering roles.

Engagements

Bring in the team

IGonics is my consultancy. Data platforms and warehousing, analytics and BI, AI and automation, systems integration and custom software — scoped, delivered and supported by a team rather than one pair of hands.