Well-Typed and the entire hs-bindgen team—there are quite a few of us, see
below—are delighted to announce hs-bindgen 1.0, the first official release.
In case you missed the announcement of the alpha release,
hs-bindgen is a tool to automatically generate Haskell bindings from C
headers.
In this blog post we will start with a very brief intro to hs-bindgen, but
we will not repeat everything that we discussed in the first two blog posts that
announced the first and second alpha releases.
Instead, we will highlight a few of the major changes; for a full list (which is
quite long!), please see the changelog.
You can find hs-bindgen on Hackage.
Introduction
Consider this small C header:
#include <stddef.h>
typedef size_t length;
typedef struct str {
char* bytes;
length len;
} str;
/**
* Construct \ref str from null-terminated input
*/
str new_string(const char* str);When we run hs-bindgen on this example
$ hs-bindgen-cli preprocess example.h --omit-field-prefixes # .. more args
it will generate a number of Haskell modules, which we summarize below:
newtype Length = Length{
unwrap :: CSize
}
data Str = Str{
bytes :: Ptr CChar
, len :: Length
}
-- .. bunch of instances ..
{-| Construct 'Str' from null-terminated input
__C declaration:__ @new_string@
.. other Haddocks elided, here and elsewhere ..
-}
new_string ::
PtrConst CChar -- ^ __C declaration:__ @str@
-> IO Str
new_string = {- .. -}Some notable features:
- C names are translated into Haskell names, modifying them to match Haskell’s naming rules where needed.
- With
--omit-field-prefixesthe Haskell code relies onDuplicateRecordFields; without that option it would have prefixed every field (lengthUnwrap,strBytes, etc.). - No Haskell definition is generated for
size_t, defined instddef.h; instead, it is mapped to the existingCSizetype by means of a binding specification for the standard C library. Binding specifications make binding generation compositional: bindings generated for one library can be reused in another. - The
lengthtypedef is translated to a Haskell newtype, but thestrtypedef around thestrstruct is not: the latter exists in C only for syntactic convenience. - Function
new_stringreturns a struct by value, which is not natively supported by Haskell’s FFI, and sohs-bindgenwill also generate the necessary C wrappers. - All Haskell declarations get Haddock comments linking them to the corresponding C declaration; if that C declaration has associated Doxygen comments, those are also translated.
In the remainder of this blog post we will highlight some of the improvements we’ve made since the alpha release. As you will see, there are a lot of details to get right!
FFI types
Consider a C function
void foo(MyType x)The obvious translation into Haskell is
foreign import {- .. -} foo_wrapper :: MyType -> IO ()However, this is valid only if GHC can see that MyType is a type supported
by the Haskell FFI; in particular, this means that GHC must be able to unwrap
that MyType newtype until it gets a base type from a fixed set of permitted
foreign types.
This is hard to ensure in general, especially if MyType is not newly
generated, but comes from another library (previously generated by hs-bindgen,
or handwritten). In the alpha release we resolved this by reducing every
argument to a base type, but this resulted in code that was less portable than
it could have been; for example, we might use Int32 instead of MyType.
We now solve this in a better way: given a Haskell type corresponding to a C
type, a binding specification can (and should) specify an “FFI
type” for that Haskell type: an FFI type is a pair of a module
M and type F such that F is legal to appear in a foreign import in a
context where M is in scope. For example, Foreign.C.CSize is an FFI type,
because CSize can appear in a foreign import if Foreign.C is imported:
Foreign.C exports all the necessary newtype constructors.
The translation from MyType to its FFI type (which might be MyType itself)
is now also more general; MyType just needs to have an instance of the
HasFFIType type class:
class HasFFIType a where
type FFIType a :: Type
toFFIType :: a -> FFIType a
fromFFIType :: FFIType a -> aThe FFIType type instance should line up with the FFI type specified in the
binding spec. In many cases, toFFIType and fromFFIType can simply be
coerce.
As an example, if MyType’s FFI type is CSize, then hs-bindgen will
generate code
foreign import {- .. -} foo_wrapper :: CSize -> IO ()
foo :: MyType -> IO ()
foo x = foo_wrapper (toFFIType x)In general the bindings generated by hs-bindgen are still specific to the
specified target platform (the host, by default), but the introduction of
FFI types reduces an avoidable source of non-portability.
Unnamed declarations
In Haskell, we might define a binary tree with values in the leaves as
data BTree a = Leaf a | Branch (BTree a) (BTree a)An analogous definition in C might look something like this:
enum tag {LEAF, BRANCH};
struct BTree {
enum tag constr;
union {
int value;
struct {
struct BTree* left;
struct BTree* right;
} branch;
};
};The union is anonymous: in C, its fields are accessed as if
they are fields of the parent (struct BTree). In addition, the inner struct
is unnamed, although the field of the union in this case does have a name
(branch).
Haskell does not have anonymous records or anonymous sums, and so we must choose
names. We have a policy in hs-bindgen to never generate numbered names
(“struct 1”), and instead pick names based on surrounding clues, to avoid the
generated code changing in unpredictable ways when the C headers change. In this
case, the union is named after its first field and its parent:
BTree_anon'value; similarly the struct is named after its parent and the name
of the field: BTree_anon'value_branch. The anon' infix avoids name clashes
with other generated names (C names cannot contain ticks).
newtype Tag = Tag CUInt
pattern LEAF, BRANCH :: Tag
data BTree = BTree{
constr :: Tag
, anon'value :: BTree_anon'value
}
newtype BTree_anon'value = ..
instance HasField "value" BTree CInt
instance HasField "branch" BTree BTree_anon'value_branch
instance HasField "value" (Ptr BTree) (Ptr CInt)
instance HasField "branch" (Ptr BTree) (Ptr BTree_anon'value_branch)
data BTree_anon'value_branch = BTree_anon'value_branch{
left :: Ptr BTree
, right :: Ptr BTree
}These names are perhaps a bit awkward, but in most cases code never has to deal
with them, because hs-bindgen generates “indirect” HasField instances that make
it possible to treat the fields of the union as if they are fields of the
parent, just like in C. Moreover, we have these instances also for pointers;
this means that when using the OverloadedRecordDot extension, we can write
Haskell code that processes these C values in a very natural way:1
toBTree :: Ptr C.BTree -> IO (BTree CInt)
toBTree bt = do
tag <- peek bt.constr
case tag of
C.LEAF ->
Leaf <$> peek bt.value
C.BRANCH -> do
left <- peek bt.branch.left
right <- peek bt.branch.right
Branch <$> toBTree left <*> toBTree rightImproved record-dot support
Although support in GHC for overloaded record update is sadly not yet as
polished as overloaded record access, it can
nonetheless be quite helpful when dealing with C code that uses a lot of unions:
unions don’t have a natural representation in Haskell, and hs-bindgen
therefore encodes them as opaque datatypes.
To continue the example from the previous section, we can construct a value of
type BTree as
import HsBindgen.Runtime.Union qualified as Union
btree1 :: C.BTree
btree1 = C.BTree{
constr = C.BRANCH
, anon'value = Union.zero{
branch = C.BTree_anon'value_branch{
left = nullPtr
, right = nullPtr
}
}
}If we take advantage of the indirect fields described in the previous section, we can improve this further:
import HsBindgen.Runtime.Struct qualified as Struct
btree2 :: C.BTree
btree2 = Struct.zero{
constr = C.BRANCH
, branch = Struct.zero{
left = nullPtr
, right = nullPtr
}
}Note that OverloadedRecordUpdate currently also requires RebindableSyntax;
hs-bindgen-runtime provides a module HsBindgen.Runtime.Overloading that
restores the standard environment. See also Enabling record dot
syntax in the manual.
Macros
Preprocessor macros are used extensively in C development, and so for many
applications it is important they too get translated to Haskell. Macro handling
in hs-bindgen has been significantly improved since the alpha. Most of those
changes are internal, though they will become user-facing once we release
hs-bindgen-as-a-library, at which point it will also be
possible to substitute your own macro languages, replacing hs-bindgen’s
default c-expr.
The full list of improvements can be found in the changelog; here we will mention just a few. First, we can now parse declarations that use a mixture of macros we can parse and macros that we cannot. For example, given
#define DeviceId int
#define MUST_CHECK __attribute__((warn_unused_result))
int upgradeFirmware(DeviceId did) MUST_CHECK;where the definition of DeviceId is parsed by hs-bindgen but the definition
of MUST_CHECK is not, we now generate
newtype DeviceId = DeviceId CInt
upgradeFirmware :: DeviceId -> IO CIntSecond, detection of ambiguous macros (macros that do not have a unique expansion throughout the translation unit) is improved. For example, given
#define A 5
#define B A
#define B A
#define X 1
#define Y X
#undef X
#define X 2
#define Y Xwe generate bindings for A and B
a :: CInt
a = 5
b :: CInt
b = abut not for X and Y.
Finally, macros can now be defined as part of the input to hs-bindgen
(--hash-define FOO 1). Previously such definitions had to be given twice, once
to hs-bindgen and once to the C compiler via ghc-options: -optc-DFOO, which
was easy to get wrong and could be hard to debug.
Conclusions
Although Haskell’s foreign function interface is good, writing bindings to
large C libraries is nonetheless a laborious process. The “fire and forget”
approach that hs-bindgen offers takes care of a lot of the legwork. We’re
pretty proud of the fact that hs-bindgen can handle almost everything that
a C header might throw at it.
However, for many applications the bindings generated by hs-bindgen will be a
starting point only: they are a direct translation of the C code, and are
therefore very low-level. We are working on a
binding-combinators library that makes this
process easier, but ultimately such high-level bindings are then still written
by hand, albeit assisted by a library. The next big goal for hs-bindgen is
automatic generation of high-level bindings, guided by pluggable heuristics.
It has taken a lot of effort to get here, and we thank
Anduril for sponsoring this work. Within Well-Typed,
the hs-bindgen team includes or at one point included Armando Santos, Dominik
Schrempf, Edsko de Vries, Finley McIlwaine, Joris Dral, Matthias Heinzel, Oleg
Grenrus, Sam Derbyshire and Travis Cardwell. We have also benefited from
numerous contributions from external contributors in the form of issues and PRs,
for which we are grateful!
hs-bindgen 1.0 is available on Hackage now; give it a spin and let us
know if you have any problems!