Python-like syntax. Native C11. ML, backend servers, systems code — in one language.
Ocean is an experimental compiled language that keeps Python-like syntax while lowering programs to native C11.
The project is designed around two first-class directions:
Ocean
├── Backend / systems
│ ├── HTTP servers
│ ├── TCP sockets
│ ├── JSON
│ ├── files
│ ├── routing / middleware
│ └── native C/POSIX interop
│
└── ML / HPC
├── Tensor
├── autograd
├── Transformer
├── TinyGPT
├── CPU / GPU
└── OpenMP
The idea is simple:
write code that feels close to Python, but keep native compilation, explicit ownership, predictable runtime behavior, and direct access to C.
Ocean aims to combine:
- Python-like syntax
- native C11 output
- static validation
- automatic ownership management
- ARC + borrowing
- C/POSIX FFI
- Tensor + autograd
- Transformer / GPT building blocks
- CPU + GPU device API
- OpenMP
- HTTP/backend runtime
A small Ocean program still looks familiar:
def square(x: float32) -> float32:
return x * x
def main() -> int:
var value: float32 = square(4.0)
print(value)
return 0
But the compiler pipeline is:
Ocean source
↓
Parser
↓
Typed IR
↓
Validator
↓
CCodeGenerator
↓
C11
↓
native executable
Ocean already has a small eager ML stack:
Tensor
Parameter
Module
Linear
ReLU
LayerNorm
MultiHeadAttention
TransformerBlock
Embedding
CrossEntropyLoss
SGD
AdamW
TinyGPT
A minimal training loop looks intentionally familiar:
import <std/tensor/tensor.oc>
import <std/ml/nn.oc>
import <std/ml/optim.oc>
def main() -> int:
var x: Tensor[float32] = Tensor.from_list(
[[0.0], [1.0], [2.0], [3.0]],
"cpu"
)
var y: Tensor[float32] = Tensor.from_list(
[[1.0], [3.0], [5.0], [7.0]],
"cpu"
)
var model: Linear = Linear(1, 1)
var optimizer: AdamW = AdamW(
model.parameters(),
0.05,
0.9,
0.999,
0.00000001,
0.01
)
var loss_fn: MSELoss = MSELoss()
var step: int = 0
while step < 200:
optimizer.zero_grad()
var prediction: Tensor[float32] = model.forward(x)
var loss: Tensor[float32] = loss_fn.forward(prediction, y)
loss.backward()
optimizer.step()
step = step + 1
return 0Under the hood this goes through Ocean's own Tensor runtime and dynamic autograd engine.
No Python runtime is required by the generated binary.
Ocean can already train a small decoder-only Transformer end-to-end.
The current TinyGPT stack is:
token embedding
+
position embedding
↓
TransformerBlock
↓
TransformerBlock
↓
LayerNorm
↓
lm_head
↓
CrossEntropyLoss
A real verified toy training run produced:
initial loss = 2.079442
final loss = 0.057157
predicted next token = 7
token embedding grad = 1
position embedding grad = 1
lm head grad = 1
That means the full path works:
Embedding
→ attention
→ residuals
→ LayerNorm
→ FFN
→ logits
→ CrossEntropy
→ backward
→ optimizer
examples/ML/gpt2_native_ternary_inference.oc measures generation of new
tokens for the canonical GPT-2 small profile (50257, 1024, 768, 12, 3072, 12). The benchmark runs only eval()/inference, disables autograd graph
construction, performs one warmup token, and prints elapsed seconds,
milliseconds per token, and tokens per second:
ocean run examples/ML/gpt2_native_ternary_inference.oc \
--cflags "-I${CONDA_PREFIX}/include -L/usr/local/cuda/targets/x86_64-linux/lib -lOpenCL -DOCEAN_TENSOR_ENABLE_OPENCL"The benchmark now uses a per-layer KV-cache: the prompt is prefetched once and
each generated token computes only the new query/key/value row. The cache uses
GPU-resident [B, H, max_seq, head_dim] tensors and native OpenCL kernels for
cache-row writes and prefix reads.
var block: TransformerBlock = TransformerBlock(
128, # d_model
8, # heads
512 # feed-forward size
)
var hidden: Tensor[float32] = block.forward(
input,
causal_mask
)The attention path supports Transformer-shaped tensors such as:
[B, H, T, D]
with batched matmul, transpose, permute, broadcasting, softmax, masking and autograd.
Ocean exposes a backend-neutral device API:
var model: TinyGPT = TinyGPT(...)
model.to("gpu")
var tokens_gpu: Tensor[int64] = tokens.to("gpu")
var mask_gpu: Tensor[float32] = mask.to("gpu")
var logits: Tensor[float32] = model.forward(
tokens_gpu,
positions_gpu,
mask_gpu
)And training keeps the same API:
model.to("gpu")
var optimizer: AdamW = AdamW(
model.parameters(),
0.001,
0.9,
0.999,
0.00000001,
0.01
)
loss.backward()
optimizer.step()The current GPU backend is based on OpenCL.
The public API intentionally stays:
"cpu"
"gpu"
instead of exposing OpenCL/CUDA-specific objects to application code.
Some Tensor operations are already GPU-native, while others still use
correctness-first CPU fallback internally. A CUDA backend can later live behind
the same "gpu" API.
OpenCL:
ocean build examples/ML/gpt2_native_ternary_inference.oc \
--cflag=-lOpenCL \
--cflag=-I"/usr/include/CL/" \
--cflag=-L"/usr/lib/x86_64-linux-gnu/libOpenCL.so" \
--cflag=-O3 \
--cflag=-DOCEAN_TENSOR_ENABLE_OPENCLCUDA:
export OCEAN_TENSOR_GPU_BACKEND=cuda
python main.py build \
./examples/ML/gpt2_native_ternary_inference.oc \
--compiler nvcc \
--cflags "-O3 -DOCEAN_TENSOR_ENABLE_CUDA -DOCEAN_TENSOR_ENABLE_OPENCL"Ocean is not only an ML language.
The standard library is also developing a native backend stack:
std/net
├── TCP sockets
├── HTTP
├── Request / Response
├── App
├── Router
├── middleware
├── worker pool
└── keep-alive
The goal is to write backend services with Python-like ergonomics and compile them into native executables.
The intended API looks like this:
import <std/net/web.oc>
import <std/json/json.oc>
def health(request: Request) -> Response:
var body: Json = Json.object()
var status: Json = Json.str("ok")
body.set("status", status)
return Response.json_value(body)
def hello(request: Request) -> Response:
var name: str = request.query_param(
"name",
"Ocean"
)
return Response.text(
"Hello, " + name + "!"
)
def main() -> int:
var app: App = App.create()
app.get("/health", health)
app.get("/hello", hello)
app.workers(8)
app.queue_size(256)
app.keep_alive(5000)
app.run("0.0.0.0", 8080)
return 0The runtime is built on C11/POSIX primitives rather than a Python event loop.
Large applications can be organized by route groups:
var app: App = App.create()
var api: Router = Router.create("/api/v1")
api.get("/users/{id}", get_user)
api.post("/users", create_user)
api.patch("/users/{id}", update_user)
api.delete("/users/{id}", delete_user)
app.include(api)Which produces routes like:
GET /api/v1/users/{id}
POST /api/v1/users
PATCH /api/v1/users/{id}
DELETE /api/v1/users/{id}
The intended middleware API is deliberately familiar:
def request_log(
request: Request,
call_next: Next
) -> Response:
print("before request")
var response: Response = call_next.call(request)
print("after request")
return response
app.middleware(request_log)The same mechanism can support:
CORS
authentication
request IDs
access logging
timing
rate limiting
error recovery
The backend runtime is designed around a bounded worker pool:
accept thread
↓
bounded connection queue
↓
worker #1
worker #2
...
worker #N
This gives predictable concurrency and memory usage without spawning an unbounded thread per request.
A particularly interesting Ocean use case is combining backend and ML in the same native program.
For example:
HTTP request
↓
JSON parsing
↓
Tensor preprocessing
↓
TinyGPT / model inference
↓
JSON response
Conceptually:
def predict(request: Request) -> Response:
var payload: Json = request.json()
var tokens: Tensor[int64] = encode(payload)
var tokens_gpu: Tensor[int64] = tokens.to("gpu")
var logits: Tensor[float32] = model.forward(
tokens_gpu,
positions,
causal_mask
)
var result: Json = decode_logits(logits)
return Response.json_value(result)That is one of the core directions of Ocean:
native backend + native ML without crossing a Python/C boundary.
Ocean also supports an ownership-safe subset of OpenMP:
#pragma omp parallel for collapse(2) schedule(static)
for i in range(rows):
for j in range(columns):
output[i, j] = left[i, j] + right[i, j]The compiler automatically adds OpenMP compiler flags when required.
Ocean separates memory behavior by type:
| Type | Runtime model |
|---|---|
int, float, bool |
plain values |
array[T] |
unique-owned buffer |
list[T], dict[K,V], classes |
ARC |
Tensor[T] |
managed Tensor object |
&T |
immutable borrow |
&mut T |
exclusive mutable borrow |
| raw pointers / C calls | explicit unsafe boundary |
A mutable borrow looks like:
def scale(
values: &mut array[float32],
factor: float32
) -> None:
for i in range(len(values)):
values[i] = values[i] * factor
return None
Borrowing does not imply a heap allocation or a reference-count increment.
Ocean can call native C APIs through an explicit unsafe boundary:
cimport <math.h>
def main() -> int:
unsafe:
var value: float64 = @sqrt(16.0)
print(value)
return 0This keeps low-level integration possible without exposing raw C semantics through ordinary safe code.
Ocean Tensor can load standard NumPy .npy files:
var weights: Tensor[float32] = Tensor.load_npy(
"weights.npy",
"cpu"
)
weights.save_npy(
"weights_copy.npy"
)This makes it possible to move model weights between Python tooling and Ocean without inventing a custom binary format.
Create an environment:
python -m venv .venv
source .venv/bin/activate
python -m pip install -e ".[dev]"Run tests:
python -m pytestBuild an Ocean program:
python main.py buildOr, after editable installation:
ocean --helpA minimal project can use:
[package]
name = "my_app"
version = "0.1.0"
entry = "src/main.oc"
source = "src"
build = "build"
[build]
compiler = "gcc"
cflags = ["-std=c11"]
[build.profiles.release]
cflags = ["-O3"]Then:
ocean check
ocean build
ocean runOcean is an active prototype.
The current verified CPU/ML regression baseline is:
138 passed
There is also a GPU integration test. With the micromamba base OpenCL setup,
both GPU integration tests run against the local NVIDIA OpenCL runtime.
Important distinction:
GPU test passed -> GPU execution verified
GPU test skipped -> environment not verified
For the micromamba base environment:
eval "$(micromamba shell hook -s bash)"
micromamba activate base
export PKG_CONFIG_PATH="$CONDA_PREFIX/lib/pkgconfig${PKG_CONFIG_PATH:+:$PKG_CONFIG_PATH}"
python -m pytest tests/test_gpu_training_v01_ocean.py tests/test_gpu_hotpaths_v01_runtime.py -qCurrent working milestones include:
Typed IR
ownership / ARC
borrowing
Tensor
ND broadcasting
batched matmul
autograd
LayerNorm
MultiHeadAttention
TransformerBlock
Embedding
CrossEntropyLoss
SGD
AdamW
TinyGPT training
OpenMP
File IO
NumPy .npy
std/net backend foundation
Near-term priorities:
GPU-native Transformer kernels
TinyGPT.generate()
KV cache
RoPE
Tensor.arange
GPU-native AdamW
CUDA backend
mixed precision
quantization
Router stabilization
nested routers
middleware
keep-alive
graceful shutdown
CORS
cookies
uploads
streaming
WebSocket
OpenAPI
database ecosystem
stronger borrow/data-flow analysis
better diagnostics
traits/interfaces
Result / Option
zero-copy Tensor views
memory pools
LSP / tooling
For architecture and implementation details:
docs/Handoff.md— current technical state and roadmapAGENTS.md— repository rules for coding agentsdocs/MEMORY_MODEL.md— ownership and memory modelstd/tensor/README.md— Tensor APIstd/tensor/OpenCL.md— GPU backend designstd/net/README.md— networking/backend stack
Ocean is exploring a space between:
Python
simplicity
C / C++
native runtime control
Rust
ownership discipline
PyTorch
ML ergonomics
FastAPI-like frameworks
backend ergonomics
The goal is not to clone any one of them.
The goal is to make code like this feel natural:
model.to("gpu")
app.post("/predict", predict)
app.run("0.0.0.0", 8080)while still ending up with a native executable.
See the repository license for details.
